GPT-6 Astra Still Can't Translate These Languages
The Most Powerful AI Model Still Fails Most Hard Translation Cases
OpenAI released GPT-6 Astra on September 3, 2026, and the announcements were loud. Nvidia's Jensen Huang declared AGI had arrived. OpenAI's Greg Brockman said the company had entered "the AGI era." The press ran with it.
Then third-party evaluators tested how well Astra actually translates.
The results, reported by Slator, tell a more complicated story. Astra leads the field on two independent benchmarks — RWS TrainAI's M-GATE across 30 languages, and the Last Translation Benchmark built from cases that have broken every major AI system. But on those hard cases, Astra's standard and Pro configurations passed only around 44–45% of them. It failed more than half. And on Finnish, it actually regressed compared to an older OpenAI model.
Tomáš Burkert, Head of Innovation at TrainAI, put it bluntly: "If AGI is here, it only really speaks a few languages and still has a lot to learn."
He's right. And the languages it speaks least well are exactly the ones that matter most to the communities TXLOC serves.
What the Benchmarks Actually Measure — and What They Ignore
M-GATE tests 30 languages. The Last Translation Benchmark tests an unspecified set of difficult cases submitted by human contributors. Neither benchmark includes Chuukese. Neither includes Pohnpeian. Neither includes Marshallese, Kosraean, or Yapese.
This is not a criticism of the researchers. Benchmark design requires verified test data, native-speaker annotators, and enough training signal to even know what correct output looks like. For low-resource Pacific languages, that infrastructure largely does not exist — which is precisely why these languages are absent.
The practical consequence: when OpenAI claims leading multilingual performance, that claim is silent on an entire region of the Pacific. There is no score for Chuukese because no one tested it. There is no regression warning for Pohnpeian because no one caught one.
Users of healthcare systems in Hawaii, Guam, and the US Pacific territories, or students in Micronesian-diaspora school districts across the continental US, don't show up in any AGI benchmark. Their language needs are invisible to the metrics being used to declare translation solved.
Why This Gap Has Real Consequences in Healthcare and Government
Consider a Chuukese-speaking patient at a federally qualified health center in Honolulu. Federal law — specifically Title VI of the Civil Rights Act and the Affordable Care Act's Section 1557 — requires that the patient receive meaningful access to care in their language. "Meaningful access" is a legal standard, not a product feature.
If a health system routes that patient's discharge instructions through a general-purpose LLM because the model topped a benchmark covering 30 other languages, that system is taking on compliance risk it may not fully understand. Astra's benchmark scores say nothing about Chuukese accuracy, fluency, or medical register. The model has no verified training data for that language pair. The output could be grammatically plausible and factually wrong in ways no one in the workflow can catch.
The same problem applies to school districts serving Micronesian-heritage families, to government agencies issuing public health notices, and to courts requiring certified interpretation. In those settings, a passing score on a benchmark you weren't included in is not a credential.
Where AI Fits — and Where It Doesn't
None of this means AI translation has no role. For high-resource language pairs with strong benchmark coverage, AI-assisted workflows genuinely reduce cost and turnaround time. The M-GATE results show real gains in languages like Tagalog, Tamil, and Brazilian Portuguese. For those pairs, a well-supervised MTPE workflow makes sense.
The breakdown comes when organizations apply high-resource assumptions to low-resource realities. Here's a practical framework:
| Language pair | Benchmark data available | AI-first workflow appropriate | Recommended approach |
|---|---|---|---|
| English — Spanish | Extensive | Yes, with post-editing | MTPE with specialist review |
| English — Tagalog | Moderate | Conditional | MTPE with native review |
| English — Chuukese | None | No | Human translation, certified |
| English — Pohnpeian | None | No | Human translation, certified |
| English — Marshallese | Minimal | No | Human translation, certified |
For the bottom three rows, there is currently no responsible shortcut. The benchmark silence is itself the answer.
The AGI Framing Is Doing Harm
When executives declare AGI has arrived and the press amplifies it, procurement teams hear something simpler: translation is a solved problem. That belief leads to budget cuts for interpreter services, pressure to drop human review, and procurement decisions that favor general-purpose tools over specialist ones.
For the communities that rely on high-resource languages, that pressure is uncomfortable. For Chuukese and Pohnpeian speakers accessing healthcare or legal services, it is dangerous.
Astra is the best AI translation model currently benchmarked. That is genuinely impressive. It is also a model that fails more than half of hard translation cases across languages it was actually tested on, and has never been publicly evaluated on the languages spoken by tens of thousands of US residents with federal language rights.
"The most intelligent model ever built" is a product claim. "Suitable for healthcare translation in Chuukese" requires evidence that does not yet exist.
The Specific Takeaway
Before routing any translation task through an LLM, check whether your target language appears in the benchmark used to evaluate that model. If it does not — and for Pacific and Micronesian languages it almost certainly does not — you are operating without a safety net in a context where errors carry legal and clinical consequences. Verified human translators with documented competency in the specific language are not a legacy fallback; they are the current standard of care.
If you need qualified Chuukese or Pohnpeian translation for a healthcare or government project, contact TXLOC to discuss your requirements.
Manages the TXLOC platform and content.
Related articles
GPT for Translation Works Until It Doesn't
The Fluency Trap You run a batch of content through GPT and the output looks clean. Readable, natural, close...
AI Optimizes the Wrong Part of Your Docs Pipeline
The Pipeline Problem Nobody Is Talking About The localization industry has spent the last three years arguing...
Why AI Translation Benchmarks Miss the Hardest Languages
A New Benchmark That Still Leaves Millions Behind Researchers have begun crowdsourcing the hardest cases in m...