As tools like ChatGPT and other generative AI systems rapidly reshape how people work, learn and communicate, most of their power is concentrated in a small group of dominant languages such as English, Chinese and French. But with more than 7,000 languages spoken worldwide, billions of people are left with limited or inaccurate AI support in their own languages.
Senior Lecturer in Te Kunenga ki Pūrehuroa Massey University's School of Mathematical and Computational Sciences Dr Surangika Ranathunga is working to change that. Originally from Sri Lanka, Dr Ranathunga has firsthand experience of the challenges speakers of low-resource languages face in the digital world.
“AI models often reflect Western ideologies, which do not reflect all languages and our cultures. The challenge is how to make AI inclusive so it supports other languages, and through that, supports other cultures,” Dr Ranathunga explains.
Alongside a team of researchers, Dr Ranathunga is working to address the problem through four connected approaches: raising awareness through position papers that highlight the digital imbalance; building datasets essential for training AI systems in underrepresented languages; developing new techniques designed specifically for low-resource languages; and evaluating existing AI systems for different tasks in the context of low-resource languages.
In addition to building datasets from scratch, her team explores synthetic data generation techniques such as web mining, alongside optical character recognition (OCR) which helps unlock information from printed materials that have never been digitised.
“AI is nothing without data. Even if people want to build tools for these languages, often the data simply doesn’t exist,” Dr Ranathunga says.
Without intervention, she warns, the AI revolution could deepen global inequality. As more of life moves into digital systems, people will naturally shift toward languages that AI supports better.
“Over time, that could have a detrimental impact and lead to the decline of underrepresented languages. Education is a key example. Even now, most tutoring systems and learning materials are designed for English. If similar tools existed for other languages, access to education would increase significantly.”
AI also struggles with tasks that require cultural context, such as generating appropriate mathematical word problems.
“We’ve seen some absurd examples, like a question where someone travels from Britain to Sri Lanka, bringing Ceylon tea as a gift. It shows the model doesn’t understand local context, and while the questions may be mathematically correct, they’re often culturally inappropriate for students in those countries.”
She explains that one of the biggest challenges is that AI tools are often considered successful once they work well in English, even though they may not work in many other languages.
“In English, spelling correction is considered solved to a great extent, as tools such as Microsoft Word spell check handles it. But in many languages, there is no spelling correction system at all. You can’t call it a solved problem until it is solved for all languages.”
Dr Ranathunga also serves on the AI advisory committee established by the government of Sri Lanka, helping guide national strategy on local language technology. One of her key goals has been developing a machine translation system for the languages used in Sri Lanka: Sinhala, Tamil and English.
“Right now, we mostly rely on tools like Google Translate, which perform poorly for low-resource languages. This also relates to the problem of AI sovereignty, our data goes to systems hosted overseas and we don’t have control of it.”
Her vision is for countries to build and host their own AI systems that serve their own languages and communities.
In the coming months, Dr Ranathunga and her team will be releasing one of the largest datasets of its kind for Sinhala-Tamil-English Machine Translation, containing more than 100,000 parallel sentences. Originally supported by a diversity and inclusion grant from Google, the project has been running for three years and is expected to have immediate national impact in Sri Lanka.
“The government of Sri Lanka is now considering building Machine Translation systems for the local languages. Instead of starting from scratch, they can build directly on the models and the datasets that we have built,” Dr Ranathunga says.
The dataset and the Machine Translation models will also enable future research in areas such as adversarial robustness, error handling and cross-lingual model improvement.
Because low-resource language research is often underfunded, as a result of many of the languages being used in developing countries, Dr Ranathunga collaborates with researchers across countries such as Sri Lanka, United Kingdom, India, Albania and Pakistan.
“This is participatory research. We come together to build datasets, evaluate models and introduce new systems. It is very time consuming, and there is not enough funding in many regions, so collaboration is essential,” she says.
Dr Ranathunga is a strong advocate for open research and hopes to see more researchers and AI specialists follow suit.
“It’s really important that we share what we build, from datasets to codes and models, because that is how this field grows. If everyone keeps their work private, there is no progress. Low-resource languages suffer, especially when resources are not shared. We need to build together.”
Related news
Using artificial intelligence to unlock how the brain recalls memories
A recent study combining neuroscience and artificial intelligence (AI) sheds new light on how memory-related regions of the brain communicate, laying important groundwork for future research into memory disorders.
New artificial intelligence major set to launch in 2026
A new major spotlighting artificial intelligence (AI) will be added to the Bachelor of Information Sciences in 2026, helping prepare tomorrow’s technology leaders for a changing world.
New webinar series to provide bite-sized science for curious minds
The College of Sciences has launched a new webinar series aimed at bringing science to the forefront of the conversation in a way that’s both accessible and engaging.