AMTA's QE Framework Misses a Critical Gap
A New QE Playbook — And One Big Blind Spot
AMTA's Quality Estimation Users Working Group released its Principles and Recipe for Evaluating Quality Estimation Systems in August 2026. It is the most structured public guidance the industry has seen on how to actually test a QE tool before trusting it with your workflows. If you buy, build, or integrate machine translation, you should read it.
But the framework has a gap that no one in the working group flagged explicitly, and it matters enormously for healthcare systems, school districts, and government agencies serving Pacific Islander communities. The guidance assumes you can assemble a representative gold-standard dataset. For Chuukese and Pohnpeian, that assumption fails almost immediately.
What the Framework Gets Right
The core argument is straightforward: QE performance is context-specific, and you cannot borrow someone else's benchmark results and apply them to your workflow. A QE system that performs well on Spanish medical documents may perform poorly on Tagalog legal text, even if both pairs are well-resourced.
The working group recommends testing 500 to 1,000 segments — roughly 5,000 to 10,000 words — per combination of content type, language pair, and use case. That is a concrete, actionable starting point. The framework also maps three distinct use cases:
| Use Case | What It Measures | Primary Output |
|---|---|---|
| Comparing AI translation systems | Relative MT quality across engines | System ranking |
| Segment-level quality scoring | Which segments need human review | Pass/fail routing decision |
| Automated error annotation | Specific error types and severity | Linguistic QA data |
Each use case demands different evaluation methods and different gold-standard data. Conflating them is exactly how organizations end up with QE deployments that look good on paper and fail in production.
The framework's treatment of LLM-as-a-judge systems is also worth attention. Two problems appear in testing: a "people-pleaser effect," where the model flags errors in clean translations, and output inconsistency across runs on identical inputs. The recommended fix for the second problem — setting temperature to zero — reduces variability but does not eliminate it. That is an honest, useful caveat that many vendors gloss over.
Where the Framework Runs Out of Road
The Recipe is built around the assumption that you can gather production data, build a gold standard, and iterate. That process works when your language pair has a functioning MT engine, a pool of qualified reviewers, and enough historical translation volume to generate representative segments.
For Chuukese (spoken by roughly 45,000 people, with significant diaspora populations in Guam, Hawaii, and the US mainland) and Pohnpeian (spoken by approximately 30,000 people), none of those conditions fully apply.
There is no commercial MT engine for Chuukese that you can evaluate with this framework. There is no open-source trained QE model for the pair. LLM-as-a-judge approaches produce outputs for Chuukese, but the framework's own "people-pleaser" warning becomes a much larger problem when the judge LLM has almost no training data in the target language. The model cannot reliably distinguish a fluent Chuukese sentence from a grammatically broken one, so its confidence scores are essentially noise.
Building a 500-segment gold standard for Chuukese medical content requires finding qualified bilingual reviewers with subject-matter knowledge, in a language community where the total number of professional translators in the US can be counted on two hands. The AMTA framework does not tell you what to do when that number is four.
What Rigorous QE Actually Looks Like for Low-Resource Pairs
This is where working with a specialist changes the evaluation entirely.
For Chuukese and Pohnpeian, QE cannot function as an automated routing layer the way it might for Spanish or Mandarin. There is no system to route around the human reviewer — the human reviewer is the system. What QE evaluation actually surfaces for these pairs is whether a given MT output is useful as a draft for a skilled post-editor, or whether it creates more work than starting from scratch.
That is a different question, and it requires a different evaluation design. Instead of segment-level pass/fail scoring, the relevant metric becomes post-editing time and error introduction rate: does the MT draft help the Chuukese translator work faster, and does it introduce errors the translator might miss under time pressure?
For healthcare clients using Chuukese interpreter services — emergency departments, federally qualified health centers, and school-based health programs serving Micronesian families in Arkansas, Oregon, and Hawaii — that distinction is not academic. A QE system that incorrectly clears a medication instruction from human review is a patient safety event. The AMTA framework's false-negative analysis applies here with much higher stakes than it does for marketing copy.
Rigorous evaluation for these pairs means small, carefully curated test sets built by the translators themselves, manual error annotation by qualified bilingual reviewers, and explicit documentation of which content types are outside the scope of any MT-assisted workflow entirely.
The Practical Takeaway
If you are a language service buyer evaluating QE for high-stakes content, the AMTA framework is the right starting point. Apply it exactly as written for your major language pairs. But before you extend any QE deployment to low-resource languages, ask your vendor a specific question: how many segments of Chuukese, or Somali, or Yup'ik do you have in your evaluation dataset?
If the answer is zero, you do not have QE coverage for that language — you have a system that will produce confidence scores without the data to validate them. Unvalidated confidence is more dangerous than acknowledged uncertainty.
If you are building or procuring QE workflows that need to cover Pacific Islander communities, we are glad to walk through what a realistic evaluation design looks like for these pairs.
Manages the TXLOC platform and content.
Related articles
Voice AI Has a Missing Language Problem
Speech AI is only as good as the voices it learned from. That is not a philosophical point — it is a data engi...
When 95% of Localization Is AI Verification—Except When It Isn't
The 95% Figure Everyone Is Quoting Florian Faes, Managing Director at Slator, made a claim recently that land...
AI Quality Estimation Fails Rare Languages
The Gap Nobody in the QE Conversation Mentions AI quality estimation is genuinely useful. The core idea, as S...