The Quiet Anxiety Behind Every Build vs. Buy Conversation

Building your own translation AI stack can look like control. Until you account for the infrastructure, dependencies, feedback systems, and accountability you are also choosing to own.


Build vs. buy Translation AI Enterprise localization
Full article

There is a kind of unease creeping into the way enterprise leaders talk about AI right now. It does not announce itself in board meetings. It surfaces in the smaller conversations. The ones over coffee at the end of a long planning session, or in the quiet pause after someone presents an AI roadmap with a little too much confidence.

The unease is not about whether AI works. That question has been settled. The unease is about what we are now depending on, and how much of it we actually control.

That is the question underneath every serious build vs. buy conversation happening in localization right now. And the question deserves a more honest answer than most of the market is giving it.

Three anxieties in one decision

When companies sit down to think through how they handle language at scale, three concerns tend to show up in the room at the same time.

The first concern is cost. Frontier model providers we have all come to rely on are operating at a loss - a typical venture capital model Their pricing reflects a land-grab phase, not a sustainable business. We have already started to see what happens when those economics shift. Uber burned through an annual budget for AI in four months. That was not a failure of planning. That is what happens when usage scales against a price you do not set.

The next moment to watch is the IPO window. When the largest model providers eventually go public, the pressure to deliver returns to shareholders will arrive on schedule. Whatever the pricing looks like then, it will not be optimized for the enterprises that built their workflows on the assumption that today's prices were stable.

In the long term, this may not be as big a concern, as Moore’s Law dictates, but it is absolutely a short-to-medium term issue.

The second concern is reliability. A frontier model that performs well today is not contractually obliged to perform the same way next quarter. Updates roll out. Behavior shifts. Outputs that used to be safe get flagged. Outputs that used to require careful prompting suddenly do not. For consumer applications, this is mostly an inconvenience. For an enterprise localization program serving regulated industries, it is something else entirely.

The engineering overhead that goes into ensuring sufficient guardrails are in place to even try to manage this variability cannot be understated.

The third concern is availability. This is the one that has hardened fastest. When the US Government restricted the new Anthropic models from being sold outside the country, it gave business operating outside the US legitimate cause for pause. While the ban was short-lived, it was not a theoretical exercise. It was a real event that exposed how thin the contracts underpinning many AI strategies actually are. If your business is in the EU, or operates across regions, the question of whether a model will still be available to you in six months is now a question worth asking out loud. That’s before we even start to talk about data sovereignty.

These three concerns do not exist in isolation. They compound. They all point to the same underlying issue: when you build on infrastructure you do not own, the things that make that infrastructure attractive today are precisely the things that can change without your consent.

The questions that should be in every evaluation

If the build vs. buy question is really a question about risk, then the evaluation needs to look different from a typical vendor RFP. Feature matrices and pricing benchmarks will tell you what the technology can do today. They will not tell you what your exposure looks like in three years.

Below are five questions worth putting in front of any provider, internal team, or technology partner involved in your translation AI strategy. They are not the only questions that matter. They are the ones that tend to get skipped.

Who actually owns the technology in the stack?

Many localization vendors do not own the AI they sell access to. They license it, wrap a workflow around it, and pass through the underlying capability with their margin layered on top. There is nothing inherently wrong with this model. It does mean that when something changes upstream, your vendor has limited ability to absorb the impact. Their roadmap is downstream of someone else's decisions.

The right version of this question is not just "do you own the technology" but "what is your dependency on third-party models, and what happens to me when those models change." A vendor who can answer that question with specifics has thought about it. A vendor who can not, has not.

What is the contractual reality when conditions change?

Most AI infrastructure agreements agreements were written for a market that has already changed. The pricing, the availability, the data terms, the regional restrictions: any of these can shift on the provider's timeline, not yours. The question is what your contract actually protects against, and what it does not.

This applies as much to direct provider relationships as to your localization vendor's relationships with their upstream providers. If you cannot trace the contractual chain, you do not know what your real exposure is. Where is the governance?

How does the system improve over time?

A localization program is a multi-year commitment. The quality you start with matters less than the quality you end with. So the question worth asking is how the system gets better over time against evolving performance metrics, and who is responsible for managing the improvement.

In a proprietary system with a closed feedback loop, every human decision becomes a training signal. Terminology decisions become enforceable rules. Improvements compound across projects. In a fragmented system, each gain is local. Each improvement has to be re-fought next quarter. In a frontier model setup, you wait for whatever the next general release brings and hope it does not regress on your use case.

The compounding effect is real and measurable. It just does not show up on a feature comparison sheet.

Where does accountability live when something goes wrong?

At scale, something always goes wrong somewhere. A term drifts. A legal disclaimer comes back slightly off. A regional team flags content that does not feel right. The question is not whether these things will happen but what your operation looks like when they do.

Specifically: can you trace the issue back to its source? Do you know who reviewed the content, when, with what guidance? Can you fix it in a way that the fix sticks across markets and content types? In a black-box system, the honest answer to all three is usually no. In a transparent system with a direct relationship between professional translators and your operation, the answers are knowable.

The relevant test is not whether problems happen. It is whether you can answer the questions about them with confidence.

What is the path forward as the technology shifts?

Translation AI is in the middle of its third major paradigm shift. Rule-based methods gave way to statistical methods. These gave way to the first wave of AI with neural models, and we’ve now shifted to generative AI Each transition rewrote what was possible and stranded the technology that came before.

Another transition is already visible on the horizon. Smaller, specialized models trained for specific tasks are starting to outperform the large models on the tasks they were never optimized to do well. Agentic AI is changing how systems integrate. Companies that lived through previous transitions are the ones most likely to make the next one without disrupting the operations that depend on them.

The sound question to ask is which “transitions” they have already navigated, and what they learned from each one. The honest version of this answer is harder to fake than the marketing version.

What the answers tend to reveal

The five questions do not predetermine an outcome. They surface what the actual exposure looks like under any given approach, which is more useful than a recommendation.

Apply them to a build-it-yourself strategy and the answers reveal an opportunity cost that rarely makes the slide deck. Twelve to eighteen months of engineering time before anything is customer-ready. Specialized infrastructure for translation memory, glossary enforcement, and human-in-the-loop feedback that has to be built and maintained separately. Engineers spending their best work on a problem that is not the company's core product, while the upstream models they depend on change underneath them.

→ Apply the questions to a multi-vendor, multi-model approach, sold as best-of-breed flexibility, and they reveal something harder to name. Every layer added is another variable nobody can fully explain. Every vendor handoff is a gap in accountability. Every integration is a place where brand voice, terminology, and approved style can silently drift. The technology slum lords of the localization industry are real, and the price they extract is paid in opacity rather than dollars.

→ Apply the questions to a frontier-model-plus-prompting approach and the answers reveal exposure to pricing, availability, and behavior changes that sit entirely outside your control. The technology may be impressive. The operational foundation is not stable.

→ Apply the same questions to a proprietary, purpose-built approach and a different picture emerges. Ownership of the technology, which means changes are absorbed rather than passed through. Contractual terms that match the multi-year reality of an enterprise localization program. A feedback loop that makes the system measurably better on your content over time. Direct accountability through the entire pipeline, with professional translators reachable through the pipeline rather than buried under a subcontracting chain. And a track record through previous technology transitions that says something about how the next one will be handled.

This is not a coincidence of how the questions are framed. It is what the questions reveal when applied honestly.

We believe proprietary technology, owned end to end and designed specifically for localization tasks, is the foundation that delivers predictability, control, and transparency at enterprise scale.

Predictability, control, and transparency

When you strip away the technology debate and ask what enterprise buyers actually need from their localization program, the answer comes down to three things.

Predictability. The ability to know, with reasonable certainty, what outputs will look like, what the program will cost, and what tomorrow will require. Not perfect foresight. Just the absence of unpleasant surprises in places where surprises are most expensive.

Control. The ability to fix things when they go wrong. To trace a quality issue back to where it actually originated. To make a change to terminology and have that change persist across to market, content types, and time. To say no when a vendor proposes a workflow that does not fit your business.

Transparency. The ability to see what is happening inside your operation. Which content was processed. Which humans were involved and when. What changed and why. How quality is trending. Where the costs and risks are concentrated.

These are not features. They are conditions. They are what makes a localization program defensible to a board, to a legal team, to a regional director who knows their market.

The five questions above are really a way of asking – in five different ways – whether a given approach can deliver on these three conditions. Most cannot. Not because the technology is not impressive, but because the conditions cannot be retrofitted onto a foundation that was not designed to support them.

What we believe, and what we will be honest about

We will be direct about our position, because it is fair to want to know.

We believe proprietary technology, owned end to end and designed specifically for localization tasks, is the foundation that delivers predictability, control, and transparency at enterprise scale. We believe the direction of travel in AI –toward vertical specialization and away from one-model-for-everything –supports this position rather than undermines it. We believe the companies that align their localization strategy with this direction now will be in a stronger position over the next three years than if they make a hasty decision on which model is most impressive in 2026.

That is not a neutral position. We are a company that builds proprietary translation AI, and our recommendation reflects what we are. The same is true of every other serious player in this space. The people who claim to take a neutral position in this conversation are usually selling something more expensive in the long run.

What we can offer instead of false neutrality is consistency. We have spent more than two decades on this specific problem. We have lived through every major technology transition in translation, from statistical methods through neural approaches and into large language models. We built systems in each generation. Lara, our proprietary language AI, is the product of that history. TranslationOS is the platform that operationalizes it. Together they are designed to make the five questions above answerable, not avoidable.

The conversation that should already be happening

If your company is planning serious international growth, or already operating across markets and feeling the strain of a system that was built for a smaller version of the business, then the build vs. buy conversation is not really optional. It is happening either by decision or by default.

The decision version is better. It tends to result in programs that are still standing in five years.

We talk a closer look into this in a quick 5-minute video, breaking down what enterprise localization actually looks like when it works, and the real cost when it doesn’t. If the questions in this article are the exact ones you're navigating right now, the video is a great next step.

An executive point of view

Build vs. Buy: What Actually Decides the Outcome

Most teams only evaluate the AI model. The actual result is decided by the hidden technology layer underneath.

Quality starts earlier than most teams think

Brand control depends on the right inputs

Feedback should improve what comes next

AN EXECUTIVE POINT OF VIEW

Build vs. Buy: What actually decides the outcome

Get the 3-minute video from Claudia Di Lorenzo on what determines localization quality before review begins: terminology, brand voice, approved translations, glossaries, style guides, and feedback loops.

Watch the video