Bad Source Text Breaks Every Translation

TA
TXLOC Admin
Platform Administrator
August 5, 2026 6 min read Translation
Cover illustration: Bad Source Text Breaks Every Translation

Your Vendor Isn't the Problem. Your Source Text Is.

You swap LSPs. You get a new translation memory, maybe a better MT engine. Six months later the same complaints come back, often using the exact same words. Field adoption is low, support tickets pile up, and someone on your team is already drafting the email to yet another vendor.

Brian Cho at MultiLingual laid out the core diagnosis clearly: the defects showing up in 20 target languages were already present in one. They were baked into the source document before any translator ever opened the file.

This is not a translation industry problem. It is a procurement blindspot, and it costs organizations real money in repeated re-translation cycles that never fix the underlying cause.

Translation Sets a Ceiling It Cannot Raise

A translator works with what they receive. If a source sentence can be read two different ways in English, it will be read two or three ways in every language derived from it. A translator who sends 200 queries on a 10,000-word manual is not being difficult. That query count is a diagnostic report on the source.

The structural issues that create this problem include:

  • Sentences that carry more than one meaning
  • Content organized in monolithic documents instead of reusable, task-oriented modules
  • Inconsistent terminology where the same component gets three different names
  • Translation memories that have drifted from what was actually shipped
  • Source files trapped in formats that were never designed for structured authoring

None of these problems are stylistic. You cannot fix them with a better editor or a smoother-sounding translation. They get locked in during authoring, long before anyone thinks about a target language.

Compliance Is an Authoring Decision, Not a Translation One

The highest-stakes version of this problem involves regulatory compliance. If your product sells into the US market, your safety content needs to conform to ANSI Z535.6 conventions and OSHA enforcement expectations. That means a clear signal-word hierarchy, correctly placed hazard statements, and an explicit relationship between hazard, consequence, and avoidance action.

None of that gets decided during translation. It gets decided when someone writes the source document. Hand a translator source text that buries a warning in body copy or uses "Danger" and "Caution" interchangeably, and there is no faithful translation that produces a compliant document. Achieving compliance in the target language means rebuilding the safety architecture from scratch. That is authoring work, and almost no localization contract is scoped or priced to cover it.

The same applies to IEC/IEEE 82079-1. A non-conformant source produces non-conformant output in every language, delivered on time and on budget.

Why This Problem Is Worse for Rare Language Pairs

For standard European language pairs, an experienced translator can often paper over minor source ambiguities. The linguistic distance is shorter, reference glossaries are large, and a senior reviewer will likely catch a misread sentence before it ships.

For Chuukese and Pohnpeian, that safety net does not exist.

These are oral-tradition languages with limited standardized written corpora. When we work on Chuukese health education materials or Pohnpeian school district communications, our translators cannot fall back on a 500,000-entry terminology database or a decade of parallel corpora. They are making judgment calls on meaning with far fewer contextual anchors than a French or Spanish translator would have.

An ambiguous source sentence in English might yield one reasonable interpretation in French. In Chuukese, the same ambiguity can split into interpretations that carry meaningfully different cultural weight, particularly in healthcare contexts where a patient's understanding of a diagnosis or treatment instruction is at stake. There is no graceful way to guess correctly when the source does not commit to a single meaning.

We have seen this directly in materials for Micronesian communities in US hospitals and school systems. A discharge instruction that vaguely says a patient should "rest" generates real problems in Chuukese translation because the word choices carry different implications about duration and activity level. When we flag this, the conversation almost always reveals that the English source was already inconsistent across different document versions. The translation query exposed an authoring problem that had been invisible for years.

For low-resource language pairs, source quality is not just important. It is the entire ballgame.

AI Amplifies the Problem, Not the Solution

The current enthusiasm around machine translation and LLMs has pushed attention further toward the engine and further away from the source. The unstated assumption is that a good enough model will smooth over a messy source document. It will not.

These tools reproduce source problems at scale. A structurally broken document now produces broken output across 30 languages in the time it used to take to get it wrong in one. The throughput improvement is real. What it amplifies depends entirely on what you feed it.

Organizations getting genuine value from AI in localization are almost always the ones whose source content was already well-engineered. The automation multiplied upstream work that had already been done. It was not a substitute for it.

What Fixing This Actually Looks Like

The practical steps are unglamorous but specific:

Problem Upstream Fix
Ambiguous sentences Author against IEC/IEEE 82079-1 from day one
Monolithic document structure Organize into reusable, task-oriented modules
Inconsistent terminology Enforce a style guide and termbase during authoring
Drifted translation memory Synchronize TM only with final, approved, shipped content
Format incompatibility Manage source in a structured authoring environment

None of this is technically complicated. It gets skipped because it happens far upstream from where the complaint eventually surfaces. By the time someone is unhappy with a Chuukese translation, three departments and two vendors removed from the original authoring decision, the structural cause is nearly invisible.

The Takeaway

Before you issue an RFP for a new LSP, pull three pages of your current source documentation and ask whether a skilled technical writer would sign off on it. If the answer is no, changing vendors will not help. Fix the source, then translate it — in English, in Spanish, and especially in the language pairs where there is no margin for ambiguity.

If your organization works with Pacific Island or Micronesian communities and you want a frank assessment of how your source materials hold up before translation begins, TXLOC is glad to take a look.

TA
TXLOC Admin
Platform Administrator

Manages the TXLOC platform and content.

Related articles

Ready to reach a global audience?

Get a free, no-obligation quote in hours. Tell us about your project and we'll handle the rest.