Your Quality Score Is Green. Should You Ship It?
The Dashboard Said Green. The Market Said Otherwise.
You launch 100 languages. MQM scores pass threshold. SLAs are met. Then Japan marketing calls the creative asset unusable, a Spanish-speaking stakeholder flags the CTA as flat, and a growth PM quietly starts shopping for a different French agency.
The scorecard still says green.
This exact scenario opened one of the most attended sessions at LocWorld55 Dublin, where Kathy Mok, Head of Localization at OpenAI, and Olga Beregovaya, VP of AI at Smartling, co-presented "Would You Ship This? Reframing Translation Quality for the AI Era." Their argument is worth unpacking carefully, because it exposes a structural flaw that affects every language program — including ones serving small, high-stakes communities that rarely appear in enterprise case studies.
What Shippability Actually Means
Shippability reframes quality review as a forward-looking launch decision rather than a backward-looking defect audit. The question shifts from "how many errors did we find?" to "would a local user trust this enough to act?"
That sounds like a minor rewording. The operational consequences are not minor.
Under the traditional model, a reviewer checks translation against a linguistic taxonomy and flags deviations. Under a shippability model, the reviewer takes local ownership of a deployment decision. Mok and Beregovaya identified four things that decision must cover:
- Meaning — Is the original intent intact?
- Market fit — Is this appropriate for this specific market?
- Surface risk — Does the content type (safety notice, legal disclosure, marketing CTA) raise the stakes?
- User trust — Would a real local user trust this enough to continue down the funnel?
The full session recap from Smartling goes deeper on how they operationalized this at OpenAI's scale.
Why Traditional Quality Models Fall Short Now
Traditional LQA was designed for a pace that no longer exists. Content now ships globally on a daily cadence. AI-first translation has become the operational default. Vendor partners are retraining workflows in real time.
In that environment, post-delivery error counting is a lagging indicator. By the time a formal LQA review confirms something was wrong, the content is already in market. The checkout flow already caused abandonment. The safety message already went unread.
The deeper problem is structural. Legacy quality models were built to find defects, not to assess whether a specific defect matters, on which surface, for which audience, at what level of risk. That distinction is exactly what AI translation speed exposes.
The Shippability Gap in Rare-Language Programs
Here is where the enterprise playbook breaks down completely — and where most agencies writing about shippability go quiet.
For major language pairs, shippability can be tested through existing infrastructure: in-country reviewers on staff, established style guides, vendor networks large enough to provide second opinions quickly. A Spanish CTA that tests poorly can be revised and retested within hours.
For Chuukese or Pohnpeian, none of that infrastructure exists at scale. The Chuukese-speaking community in the United States is concentrated primarily in Hawaii, Guam, and specific zip codes in states like Arkansas and Washington. There is no large pool of in-country reviewers available on demand. There is no established corpus of localized healthcare or government content to benchmark against.
This creates a shippability problem with dimensions that a standard four-point review rubric does not account for:
| Factor | Major Language (e.g., Spanish) | Rare Language (e.g., Chuukese) |
|---|---|---|
| Reviewer availability | High; multiple options | Very limited; community-based |
| Style guide / corpus | Extensive | Sparse or nonexistent |
| Community feedback loop | Fast; large user base | Slow; requires trust relationships |
| Regulatory signoff timeline | Predictable | Variable; depends on agency liaisons |
| Dialect variation risk | Manageable with regional flagging | High; Chuukese has significant island-group variation |
When a Hawaii Department of Health needs a Chuukese-language patient consent form to be shippable, the forward-looking question is not just "is the meaning intact?" It is: which dialect region does this patient population come from? Has a community-trusted reviewer seen this, not just a credentialed linguist? Does the terminology align with what community health workers actually say in clinic? Will the state's language access officer accept this for compliance purposes, and how long does that signoff take?
Those questions do not fit neatly into an MQM rubric. They require an agency that has actually built the community relationships, not one that added the language pair to a dropdown menu.
What Good Shippability Review Looks Like in Practice
For rare Pacific and Micronesian languages, a credible shippability process looks different from the enterprise model in three concrete ways.
Community reviewer involvement is non-negotiable. A linguist with credentials is necessary but not sufficient. The content needs to pass a trust test with someone who represents the actual receiving community. For Pohnpeian health materials, that means a reviewer connected to the Federated States of Micronesia diaspora, not just a person who holds a degree in the language.
Surface risk must be calibrated higher. When a Chuukese-speaking parent reads a school district's enrollment form, the consequences of a mistranslated deadline or an ambiguous consent clause are immediate and concrete. The error tolerance is functionally zero. Shippability review for these documents should reflect that, not default to the same threshold used for a marketing email.
Regulatory timelines must be built into project plans from day one. Some US healthcare and government agencies require formal language access review before materials can be distributed. For rare languages, that review process can take weeks. A shippability framework that does not account for institutional approval cycles will produce technically ready content that cannot legally ship on time.
The One Question That Changes Everything
The shift Mok and Beregovaya introduced at LocWorld55 is genuinely useful, and it is long overdue. Reorienting quality review around "would a local user trust this enough to act?" forces programs to confront the gap between linguistic correctness and real-world effectiveness.
For organizations serving Chuukese, Pohnpeian, Marshallese, or other Pacific Island communities, that question has to go further: who, specifically, represents that local user, and have you actually talked to them?
If your current quality framework cannot answer that, your green dashboard may be misleading you.
If you are building a language access program that includes Pacific Island or Micronesian languages, we are glad to walk through what a realistic shippability review process looks like for those communities.
Manages the TXLOC platform and content.
Related articles
Cooperative Contracts Are Reshaping Language Access Procurement
Global Interpreting Network just made it easier for thousands of public agencies, school districts, and health...
What HSA Got Right About Healthcare Interpreting
A Small Health System Just Set a High Bar The Health Services Authority in the Cayman Islands did something m...
Language Access for All Act: What It Means for LEP Services
A Federal Bill to Codify What an Executive Order Once Promised Senators Andy Kim (D-NJ) and Mazie Hirono (D-H...