The promise and pitfalls of artificial intelligence in Indigenous language translation

Comprehensive translation work plays a critical role in giving people access to education, resources, and related materials in their own language.
As in many language-based domains, generative artificial intelligence (AI) and machine translation (MT) have much to offer the translation industry by assuming many of the more repetitive and time-consuming tasks, and by ensuring the use of consistent terminology across substantial volumes of text. As a result of improved capacity and efficiency, translators can, in many cases, find themselves able to engage in larger projects while working within stricter deadlines.
At the same time, the downsides to using AI and MT in translation and localization are undeniable and important to consider. In many languages, especially widely-spoken ones, AI and MT can perform word-for-word translations reasonably well, but—at least as of the writing of this article—lack the necessary cultural and contextual nuance required for meaningful localization. The effects of this can be significant should translators and companies become overly reliant on their use, or neglect to check outputs for errors or cultural appropriateness.
While the management of these issues has become a normal part of many translation workflows, there are additional and specific challenges involved with the use of AI and MT to translate into or from languages with smaller speaker populations, and especially those with less representation, few legal rights, or sparse written materials. For many of these less widely-spoken languages, the challenges remain significant enough that the use of AI and MT remains strongly inadvisable.
Indigenous language translation in Canada: Barriers to Access
This article discusses Indigenous languages spoken within Canada’s borders, though many of the challenges described here are also relevant to Indigenous languages elsewhere, as well as many regional minority languages and marginalized communities across the globe.
In Canadian contexts, most projects that focus on Indigenous languages translate materials from an official language—French or English—into an Indigenous language. Much of this work occurs at the federal and provincial levels of government to help provide access to healthcare, the justice system, social services, and education. Additional translation work takes place in natural resource industries, such as mining, fishing, and forestry, that often take place on or around lands and waters legally owned by Indigenous communities. There is also a growing demand for Indigenous translation work in museums, institutions, universities, and across the public domain.
An important factor to consider looking at the translation landscape is the sheer diversity of Indigenous languages in Canada—between 70 and 90 languages spread across 12 distinct language families, each of which is entirely unrelated to the next.
Some, such as many variations of Cree and Inuktitut, have large speaker populations and enjoy vigorous use in the community. However, approximately 75% of the Indigenous languages in Canada have significantly fewer speakers, being classified by UNESCO as either endangered or on the verge of becoming dormant.
Despite an increasing demand, services for many Indigenous languages can be difficult to access due to a lack of available translators, and those competent in the language are often overburdened with existing projects. AI and MT offer exciting opportunities to build capacity by removing some of the burden from existing Indigenous translators, which is important at a time when many communities and organizations are pushing to increase the number of materials available in their languages to support speakers and learners. In their current state however, AI and MT remain incapable of providing outputs that approach—let alone rival—the quality of human translations.
A lack of available training data
While the precise amount of existing text that can be used as training data varies depending on the language, Indigenous languages have been recorded and handed down through countless generations via oral literacy and storytelling. Coupled with relatively small speaker populations and difficulties reaching a consensus on what many written standards should look like, most Indigenous languages have insufficiently large amounts of written corpora with which to train reliable language models.
Because of this, models trained with textual data in an Indigenous language risk being overly inflexible and restrictive. This is especially relevant given that Indigenous languages harness intricate morphological systems to create words that are deeply nuanced, specific, and rich in meaning. As a result, translations first performed by AI or MT that are then reviewed and vetted by experts may still fail to convey meaning effectively, because they are unable to access the full creativity and complexity these languages are naturally capable of.
The importance of lived experience and cultural knowledge
Many terms and concepts in Indigenous languages do not have direct counterparts in European languages, and require the careful attention of those with lived experience in both worlds to map ideas over in meaningful and insightful ways. What may be represented in one language with a single word requires multiple words or a sentence in the other, necessitating a depth of meta-cultural understanding that, as far as we know, only humans are capable of.
Knowing not just what to convey but how to convey it remains a pressing issue for AI- and MT-driven translations across the world’s languages, and there are additional challenges when working with Indigenous communities. This is in no small part because communities’ diverse worldviews are specific to the lands they come from, and have unique conventions and registers that extend beyond simple word formation and conjugation that can be difficult for outsiders to grasp. Think of how your own language use changes when you talk to a colleague versus a close relative, and if you are multilingual, how your approach differs in the other languages you speak.
As AI models are not known to be capable of understanding this nuance, their outputs risk offending readers, failing to capture necessary contextual subtleties, or being completely unintelligible altogether. And this is where a deeper risk arises:
If inaccurate AI outputs begin to comprise a significant part of a language’s limited written corpus, they will subtly come to form the basis for educational materials, written media, and subsequent training data, resulting in irreparable harm to that language.
Respecting community protocols and data sovereignty
AI systems are built and trained on data, and those data need to come from somewhere. Depending on the model, data may or may not be gathered with the consent of its creators, and this variability and uncertainty in how different models use data speaks to the idea that AI development and associated risk has been outpacing meaningful governance.
While existing applicable governance frameworks such as OCAP® and CARE provide infrastructure that supports Indigenous communities in deciding how their knowledge is used, they were not originally designed with AI in mind, and have not been codified into international copyright laws. As such, large-scale commercial AI developers are not legally obliged to respect Indigenous governance frameworks if they don’t suit their needs.
Any community knowledge accessible to the public online may be used as training data for AI and large language model (LLM) translations, and this information is not only difficult to verify en masse, but in most cases has not been uploaded with the intended purpose of being used as training materials in the first place.
The irreplaceability of human language experts
As of the writing of this article, there is simply not enough available textual information from virtually any Indigenous language in Canada to be able to train AI or MT models sufficiently to perform effective and complete translations. Of the data that is available, it is difficult to determine what is representative of the language as it is actually spoken, as well as what proportion has been gathered via approaches that respect data sovereignty.
These above-mentioned challenges, while non-exhaustive, make the responsible use of AI and MT in Indigenous language translation functionally impossible, at the very least for the time being. Importantly, they drive home the irreplaceability of speakers of Indigenous languages in the translation process, and the fact that the cultural and linguistic knowledge they hold is critical.
Frequently Asked Questions
Why can’t AI effectively translate Indigenous languages?
AI models lack the cultural and contextual nuance needed for meaningful translation. Indigenous languages have intricate grammatical systems that create deeply specific meanings that word-for-word AI models cannot interpret fully.
What happens if AI translation errors are not caught?
If inaccurate AI outputs are published, they risk offending readers or being completely unintelligible. Over time, these errors can subtly become part of the language’s limited written materials, causing harm to educational resources and changing the language being learned.
Why are human language experts irreplaceable in this process?
Translating between Indigenous and European languages requires lived experience in both worlds. Human experts possess the deep meta-cultural understanding needed to map complex ideas that do not have direct, single-word counterparts.
What is Indigenous data sovereignty and how does it relate to AI?
Data sovereignty is the right of Indigenous communities to govern how their knowledge is used. Current AI models often scrape public online data without consent, outpacing existing governance frameworks like OCAP® and CARE.
