As artificial intelligence becomes increasingly embedded in everyday life, a critical challenge remains largely unresolved across Africa: most African languages remain significantly underrepresented in the datasets used to train AI systems.
For Fatima Yusuf Tomsu, founder and CEO of FYT Localisation, that gap represents both a technological challenge and an economic opportunity.
Speaking with AIBase.ng reporter Ahmad Ibrahim, Tomsu said FYT Localisation is working to build the language infrastructure needed to ensure African languages, cultures, and communities are properly represented in the next generation of AI systems.
“FYT Localisation is a language services agency that deals with the processes and systems used to build AI models today,” Tomsu said. “We are building an AI and language infrastructure that focuses on spoken languages, as well as the language and cultural context surrounding them.”
According to her, the company’s broader mission is to help close the gap between rapidly advancing AI technologies and the many languages that major technology platforms have overlooked.
Why African languages matter for AI
The rise of generative AI has created unprecedented demand for high-quality language data. However, while languages such as English, Spanish, and Mandarin are heavily represented in AI training datasets, many African languages remain largely absent.
Tomsu believes this imbalance threatens to leave millions of Africans underserved by emerging technologies.
“The main issue is that when you look at how smart devices and technology are being used, people in local communities are increasingly using them, but there is still a significant language gap,” she told AIBase.ng.
She noted that many Africans either do not speak English fluently or prefer communicating in their native languages, even when they are highly educated.
“The problem is that many native African languages are not adequately represented in digital systems,” she said. “This is one of the things hindering a lot of development within Africa.”
According to Tomsu, the issue extends beyond simply collecting more data.
“No matter how much a model is trained, if the underlying data is not of high quality, the model will not produce high-quality output,” she said.
Building datasets for underrepresented languages
FYT Localisation has shifted its focus toward AI development services, particularly in language data collection, evaluation, annotation, transcription, and other activities that support AI training.
The company works with both structured and unstructured data, ranging from written texts and speech recordings to images, videos, and audio datasets.
“For structured data, this can include written-form data, speech, transcription, translation, annotation and other forms of language data,” Tomsu explained. “For unstructured data, we also work with different forms of media, including images, videos and audio.”
She stressed that effective AI development requires more than simply gathering information.
“It is not simply about collecting data; the data has to be structured and processed in a way that makes it useful for AI systems,” she said.
The complexity of African languages
One of the reasons many AI systems struggle with African languages, according to Tomsu, is the deep cultural and contextual complexity embedded within them.
“African languages are not like many other languages,” she said. “There are languages in Africa that have different forms, dialects and meanings depending on the context.”
She explained that literal translation often fails because many words derive meaning from cultural context, sentence structure, and even how they are spoken.
“For you to train your AI model properly, you have to make sure that the system understands what words mean in different contexts and sentences, as well as the cultural aspects associated with them,” she said.
As a result, she estimates that more than 70 per cent of African languages remain underrepresented in modern AI systems.
“There are only a limited number of African languages that are well represented in major technology platforms,” she said. “Many African languages are not even available on platforms such as Google Translate.”
Collecting data directly from communities
To build language datasets, FYT Localisation often works directly with local communities, including elders and fluent native speakers.
“You cannot always get a professional in the field to provide the language data,” Tomsu said. “Sometimes you have to go into communities and work with community members, including elders and people who are fluent in the language.”
The company has participated in projects involving language data collection in communities where languages are poorly documented and rarely represented in digital systems.
Related
- Why AI Struggles With Nigerian Names, Places, And Contexts
- How Synthetic Data as a Service Is Driving the Future of AI Models
Tomsu emphasised that ethical data collection practices are central to the company’s work.
“We have to clearly explain what we want from them and how their data will be processed and used,” she said.
According to her, contributors receive clear information regarding how their recordings and language data will be used, including written agreements outlining the purpose of collection and processing.
“The important thing is that people should understand what they are contributing to and how their data will be processed,” she said.
Why humans still matter in AI
Despite the rapid progress of generative AI systems, Tomsu argues that human expertise remains indispensable in language-related work.
“We still need the human touch,” she said.
While AI can accelerate translation and process large volumes of content, she believes cultural interpretation and contextual understanding remain uniquely human strengths.
“Translation is not just about words. It is about meaning, context and trust,” she said.
Tomsu warned that relying entirely on AI-generated translations without human review can create serious inaccuracies.
“If you simply give something to AI to translate and then give that translation to people without reviewing it, it may completely miss the meaning,” she said. “It may confuse people rather than help them understand.”
Instead, she advocates for a human-in-the-loop approach in which AI enhances efficiency while people remain responsible for final evaluation and quality assurance.
Global demand for African-language data
One of the more revealing insights from Tomsu’s interview concerns who is currently investing in African-language datasets.
According to her, much of the demand is coming from outside the continent.
“I have worked with organisations and clients outside Africa that need African-language data to build models,” she said.
She cited engagements with organisations from China, India, and the United States seeking African-language datasets for AI development.
“That tells you something about the gap we have,” she said.
According to Tomsu, African technology companies, investors, policymakers, and funding organisations need to become more involved in building the data foundations required for AI systems that accurately represent African realities.
African languages as a competitive advantage
Beyond technology, Tomsu sees African-language AI as a major commercial opportunity.
“It is very high,” she said when asked about the commercial value of African languages.
She believes growing international interest in Africa’s technology, agriculture, mining, and consumer markets will increase demand for language infrastructure that enables businesses to communicate effectively with local communities.
“Foreign companies and investors therefore have to adapt to African languages if they want to tap into the opportunities and resources available on the continent,” she said.
Building Africa’s AI future
Looking ahead, Tomsu remains optimistic about the future of African-language AI, though she acknowledges that progress will take time.
“It is going to be great. It is not going to happen overnight, but it will be possible,” she said.
She believes the next phase of AI development must be driven not only by technology companies but also by communities whose languages and cultures are being represented.
“We need to build systems that understand our languages and our realities,” she said. “If we can do that, technology will become more useful and more accessible to people.”
For Tomsu, the message is clear: African-language AI should not be left solely to organisations outside the continent.
“If we want technology to represent us properly, we have to participate in building it,” she told AIBase.ng. “Our languages, cultures and communities have to be part of the data and development process.”
