Why generative AI underperforms in non-English languages

Lost in Translation

Conversational AI is dominated by English, with serious consequences for other languages that are structural and can only be resolved with significant effort.

Ask ChatGPT a complex question in English and you’ll quite often get a correct, well-formulated answer that fits the context. Try the same in Hindi, Bengali or Yoruba, and the response tends to get shorter, vaguer, and occasionally plain wrong. In German, French or Spanish, the answers are more on-point, but they often still fall short of the content and linguistic quality of an English response.

Generative AI has a language problem. And it doesn’t only affect rare or endangered languages – it hits widely spoken ones too. The Brookings Institution describes the quality gap as a continuum: from English through European languages like German, French and Spanish, all the way to the roughly 7,000 languages spoken worldwide, of which only about 20 are considered “data-rich” – with the gap widening dramatically as you move down the list. This is a problem that surfaces repeatedly in GenAI projects. Non-English systems struggle with precision, hallucinate more frequently, and simply fabricate content that doesn’t exist.

If we have language technology that doesn’t work for people in the language that they speak, those communities don’t see the technology boost that other people might have.

Sanmi Koyejo, Stanford University, 2025

The core problem: these systems were built in English

This isn’t a bug that someone forgot to fix. It’s a structural issue baked into the architecture of virtually every language model out there. Models learn from data – and the data is overwhelmingly English. Of the content in the Common Crawl dataset, the backbone of most large language model training, over 40% is in English, while no other language comes close to 7%. Models know what they’ve seen, and most of it is in English.

Language Common Crawl Share Total Speakers % of World Population Web to Speakers Ratio
English 41.06 % ~1.53 billion ~18.7 % 2.2x
German 5.98 % ~135 million ~1.6 % 3.7x
Chinese 4.99 % ~1.18 billion ~14.4 % 0.35x
Spanish 4.66 % ~560 million ~6.8 % 0.7x
French 4.61 % ~310 million ~3.8 % 1.2x
Italian 2.38 % ~85 million ~1.0 % 2.4x
Hindi 0.22 % ~610 million ~7.4 % 0.03x

Sources: https://commoncrawl.github.io/cc-crawl-statistics/plots/languages (accessed March 30, 2026). Ethnologue 2025 (Eberhard, Simons & Fennig, eds., Ethnologue: Languages of the World, 27th ed., SIL International) – for total speaker counts (L1+L2).

The practical consequence: complex queries in non-English contexts get answered less precisely – especially in specialist domains like law or public administration. And even though German, French, Spanish and Italian are comparatively privileged as relatively “data-rich” languages, they still share the same structural problems – just in a milder form. There’s a secondary effect worth flagging too: content moderation instructions – for filtering hate speech or detecting statements that indicate serious mental health risk – are primarily designed and trained in English. In other languages, that precision degrades. Things get missed. Or get flagged when they shouldn’t be.

These issues get amplified with smaller models – what the industry calls Small and Medium Language Models. These are well-suited for RAG-based standalone deployments, which sidestep the data privacy headaches of cloud solutions. But in non-English, they’re even clumsier than their larger counterparts. There are some targeted fixes – like the embeddings from Berlin-based Jina for the small Gemma language models – but none of them fully solve the underlying problem. Even Mistral – Europe’s own French-built LLM alternative – performs better in English than in its intended target languages, German and French.

Summary: Non-English takes more work

The Brookings Institution used a quote to open its 2024 analysis of the AI language gap – and it still fits:

The limits of my language mean the limits of my world.

Ludwig Wittgenstein (1889–1951), Philosopher

The language you work in demonstrably shapes what AI can do for you. For people who don’t speak English, this translates to a material disadvantage in a world where conversational AI is becoming a useful tool for solving problems and creating output.

From a business perspective, it shouldn’t come as a surprise, then, that non-English conversational AI projects routinely underperform and often stumble already at the prototype stage. What is consistently underestimated is the additional effort involved: foundational model training for the application’s specific language patterns, a well-designed RAG pipeline, and careful fine-tuning of the language generation. All of that makes good AI implementation more expensive in non-English environments.

It takes time, money, and a deliberate commitment to linguistic diversity in the development process. But the first step is simply acknowledging the gap exists. Ignore it, and the price is an AI application that delivers little – or worse, negative – value.

References and further reading

Data Sources

Common Crawl. (2026). Crawl statistics: Distribution of languages. commoncrawl.github.io/cc-crawl-statistics/plots/languages.

Eberhard, D. M., Simons, G. F., & Fennig, C. D. (Eds.). (2025). Ethnologue: Languages of the world (27th ed.). SIL International.

Research

Alhanai, T., Kasumovic, A., Ghassemi, M., Zitzelberger, A., Lundin, J., & Chabot-Couture, G. (2025). Bridging the gap: Enhancing LLM performance for low-resource African languages with new benchmarks, fine-tuning, and cultural adjustments. Proceedings of AAAI 2025. arXiv:2412.12417.

Stanford HAI, The Asia Foundation, & University of Pretoria. (2025, April 22). Mind the (language) gap: Mapping the challenges of LLM development in low-resource language contexts. Stanford Institute for Human-Centered Artificial Intelligence white paper.

Lynch, S. (2025, May 19). How AI is leaving non-English speakers behind. Interview with Sanmi Koyejo. Stanford Report.

Yang et al. (2024). Problematic tokens: Tokenizer bias in large language models. IEEE Special Session on Privacy and Security of Big Data (PSBD 2024). arXiv:2406.11214.

Primary Sources

Wittgenstein, L. (1922). Tractatus logico-philosophicus, Proposition 5.6.