Voice AI Has a Missing Language Problem
Speech AI is only as good as the voices it learned from. That is not a philosophical point — it is a data engineering problem, and right now the data is badly skewed.
A recent Slator conversation with NCSpeech co-founders Dmitrii Sandzhiev and Iurii Agafonov put numbers and context around something localization professionals already sense: the languages that get the least AI investment are the ones spoken by people who most need reliable language access. NCSpeech is trying to close that gap by turning idle time inside super apps into paid data-collection tasks for AI labs. Drivers, riders, and passengers complete audio recording jobs and get rewarded. The recorded speech becomes training data.
It is a smart model. But it still reaches the languages that have smartphone penetration, super app ecosystems, and enough speakers to make the economics work. Chuukese and Pohnpeian are not on that list.
Why Low-Resource Means More Than Just "Rare"
When AI researchers say a language is low-resource, they mean there is not enough recorded, labeled, transcribed audio to train a usable model. The NCSpeech team flags Malaysian code-switching — speakers flipping between Malay, English, Mandarin, and local dialects — as a hard problem. That is hard. But Malaysian languages collectively have millions of speakers, growing tech infrastructure, and commercial interest from regional AI labs.
Chuukese has roughly 45,000 speakers. Pohnpeian has around 30,000. Both are spoken across the Federated States of Micronesia and inside US communities in Guam, Hawaii, and the Pacific Northwest. Neither language has a usable speech recognition model. Neither appears in any major commercial voice AI product. If a Chuukese-speaking patient calls a hospital advice line that uses voice AI triage, the system will not understand them. Full stop.
That is not an edge case in communities where Chuukese and Pohnpeian are primary languages. It is the everyday reality.
The Training Data Problem, Specifically
Building a speech model requires more than collecting audio. The NCSpeech team describes the actual pipeline: controlling recording environments, verifying speaker consistency across sessions, detecting synthetic or manipulated submissions, handling local data storage rules, and processing high volumes of files into clean, labeled datasets. Their Kazakh model now outperforms open alternatives, which tells you what properly engineered data can produce even for a language that was, until recently, deeply underserved.
For Chuukese and Pohnpeian, none of that pipeline exists yet. There is no labeled corpus. There is no benchmark model to beat. The infrastructure has to be built from scratch, which means the first requirement is not technology — it is people.
Specifically, it is trusted bilingual speakers who can produce audio, validate transcriptions, and flag dialectal variation. Chuukese, for example, has regional variation across the islands that a non-native speaker or an automated system cannot reliably detect.
What a Real Speaker Network Actually Looks Like
This is where the TXLOC perspective differs from what a generalist agency or an AI data platform can offer.
We work with a vetted network of Chuukese and Pohnpeian linguists and community interpreters — people who already operate in healthcare, legal, and education contexts. They understand register differences. They know when a speaker is using formal Pohnpeian versus the everyday spoken form. They can identify when a recording has dialectal features that would confuse a model trained only on one island's speakers.
That kind of human infrastructure is exactly what NCSpeech is trying to replicate at scale in emerging markets. For Micronesian languages, it cannot be crowdsourced through a super app. The speaker population is too small and too geographically scattered. It has to come from deliberate community relationships built over time.
The table below illustrates why the data collection approach has to differ by language context:
| Factor | High-volume approach (e.g., Malaysian dialects) | Small-community approach (e.g., Chuukese) |
|---|---|---|
| Speaker pool size | Millions | ~45,000 globally |
| Data collection method | App-based crowdsourcing | Vetted community networks |
| Dialect validation | Automated flagging feasible | Requires human expert review |
| Commercial AI interest | Growing | Minimal without institutional push |
| Primary use case | Consumer tech, voice assistants | Healthcare, legal, education access |
The Institutional Gap That Actually Matters
NCSpeech's roadmap includes expanding to more countries, building reusable dataset libraries, and developing proprietary models. That is the right direction for the languages they serve. But the institutions most responsible for Chuukese and Pohnpeian speaker communities — US healthcare systems, Pacific Island school districts, federal agencies administering Compact of Free Association obligations — are not coordinating on AI data development. They are barely coordinating on basic interpreter access.
Until a hospital system in Guam, a school district in Hawaii, or a federal health agency decides that Chuukese speech data is worth funding, the commercial market will not fill the gap on its own. The ROI calculation does not work at 45,000 speakers unless there is an institutional buyer willing to absorb the upfront data collection cost in exchange for a functional tool.
The NCSpeech story is useful precisely because it shows what is possible when someone actually does the engineering work. Their Kazakh model outperforming open alternatives is proof that low-resource does not mean impossible — it means under-invested.
The Specific Takeaway
If you work in healthcare administration, school district leadership, or a government agency serving Pacific Islander communities, the relevant question is not whether AI will eventually support Chuukese or Pohnpeian. It will, once someone pays to build the data infrastructure. The question is whether your organization will be part of funding that, or whether you will keep relying on human interpreters alone for another decade while the gap widens.
Human interpreters are not going away — nor should they. But speech AI that can handle language routing, preliminary triage flagging, or automated transcription review would reduce the burden on a very small pool of qualified Chuukese and Pohnpeian interpreters who are already stretched thin.
If you are a language service company looking for a partner with existing Chuukese or Pohnpeian linguist relationships, TXLOC is one of the few agencies in a position to have that conversation seriously.
Manages the TXLOC platform and content.
Related articles
AMTA's QE Framework Misses a Critical Gap
A New QE Playbook — And One Big Blind Spot AMTA's Quality Estimation Users Working Group released its Pr...
When 95% of Localization Is AI Verification—Except When It Isn't
The 95% Figure Everyone Is Quoting Florian Faes, Managing Director at Slator, made a claim recently that land...
AI Quality Estimation Fails Rare Languages
The Gap Nobody in the QE Conversation Mentions AI quality estimation is genuinely useful. The core idea, as S...