In Latin America and the Caribbean, millions of people face barriers to accessing financial, health, and educational services for a reason that rarely comes up in the technology debate: many digital tools do not work in their native language.
For speakers of Indigenous languages, this means they may be unable to access information, complete administrative procedures, or interact with digital services on an equal terms. This is a silent digital language gap that has profound implications for economic and social inclusion.
The scale of the challenge is considerable. According to the UNESCO, less than 2% online content is available in Indigenous languages. At the same time, 38.4% of the 556 Indigenous languages in our region are at risk of extinction, representing an 18% increase since 2009. Today, generative artificial intelligence (AI) offers an opportunity to expand access to services and knowledge in languages that have historically been excluded from the digital environment. An IDB Lab project focused on Guaraní in Paraguay illustrates both the problem and a possible solution.
AI Performance in Indigenous Languages
A study produced by IDB Lab, Microsoft, and LLYC as part of the fAIr LAC initiative, evaluated the performance of the leading large language models (LLMs), including both proprietary and open-source, in five Indigenous languages of the Americas: Quechua, Guaraní, Aymara, Nahuatl, and Quiché Maya. The results were compared with those of other, more digitially represented European languages, such as Catalan and Basque.
The findings reveal significant performance gaps. While Catalan scored 8.58 out of 10, the Indigenous languages achieved significantly lower results. Quechua performed the best, with 3.72 points, while Quiché Maya scored only 1.25.
Beyond the scores, the models struggled to consistently generate responses in the same language. Thirty-five percent of queries made in Indigenous languages received responses in another language, and an additional 11% produced responses with incoherent or nonsensical terms.
The depth and quality of the responses were also limited. On average, responses generated in Indigenous languages were four times shorter than those produced in Spanish. Furthermore, comprehension and abstraction capabilities showed poor results, while the analysis identified persistent cultural biases associated with poverty, rurality, or folkloric representations.
Why Doesn’t AI Speak an Indigenous Language Well?
There is an 84% direct correlation between the volume of open digital data available in a language and the technical performance of the models that facilitate it. LLMs do not use methodologies specific to each language; rather, they learn from the critical mass of data absorbed during their training.
In this regard, the importance of tools like Wikipedia is decisive. The number of entries in a language is the strongest predictor of the quality with which AI expresses itself (91% correlation). Added to this is a gap between model types, where proprietary models are 2.2 times more effective than open-weight models (those whose trained parameters are published for anyone to download) in Indigenous languages (4.35 versus 2.78 in Quechua) and the scarcity of processing tools: the lack of machine translators, language detectors, and speech-to-text converters hinders large-scale training.
The “GuaranIA” project in Paraguay
Overcoming this challenge requires more than just a standard translator. It requires urgently adapting technology to the people who need it most. The GuaranIA project, launched and co-financed by IDB Lab in Paraguay, is a pioneering model for addressing this challenge at its root.
The impact of this project is defined by several key milestones:
- Building the Largest Unified Digital Corpus in Guaraní. Since this language established its formal writing system in the mid-20th century and has historically relied on oral transmission, digital data was scarce. GuaranIA created the language’s largest repository, featuring texts, bibliographies, and transcribed audio recordings, all validated by the speaking communities.
- Testing Data Through Hackathons. In these events, innovators, entrepreneurs, and academics validate data to train virtual assistants, translators, and service systems that accurately interpret native cultural contexts.
- Launching Community-Based Pilot Projects.The project is developing pilot applications in three areas: health, with Guaraní-language triage via telemedicine; intercultural education, tutoring and bilingual assistants; and financial services, with adapted procedures, subsidies, and credit options.
Recommendations for a Future of Technological Inclusion
To scale these solutions, the study presents several strategic lines of action, including the promotion of enabling tools, such as voice-to-text technologies that capture oral knowledge and transform it into training data. It also addresses dialectal convergence and standardization, through agreements between language academies and institutes to establish flexible standards for Guaraní without eliminating local variants.
It is also important to build partnerships that help adapt digital tools, including operating systems, search engines, and ATMs, to local languages through conversational solutions. It is also necessary to promote training and content creation within communities so that more Indigenous youth and local creators can share knowledge and generate content on the internet and digital platforms.
Building powerful, inclusive, and culturally relevant AI is essential for Latin America’s development. Integrating Indigenous languages into the core of the digital transformation not only preserves the region’s cultural heritage but also ensures that tomorrow’s economy leaves no one behind. Having AI speak your language is more than a technological issue: it is a decision about who can access the opportunities of the digital world and be part of the future we are building.
You can read the full study here.