Will Sovereignty Policies End Africa's Language Data Flaring?
30 July 2026
Data Ecosystem and Policy
mansurat
Digital SovereigntyLanguage Data FlaringNLP AfricaAU AI StrategyLINGUA AfricaTech Policy
Will Sovereignty Policies End Africa's Language Data Flaring?
As global AI systems extract and 'flare' African language data, the continent loses critical digital power. We analyze whether the 2026 AU Continental AI Strategy and local sovereign computing infrastructure can successfully protect our linguistic heritage from foreign exploitation.
In the mid-19th century, gas flaring began in places like Baku (Azerbaijan) and Pennsylvania (U.S.) when the interest was in getting liquid petroleum for refining into kerosene for lamps, and later gasoline for cars. The natural gas that came rushing up during exploration was viewed as needless and it made sense.
In the late 20th century, the roots of language data flaring began to sprout as global computing and the internet grew based on written text infrastructure.
As defined and contextualized by Dr. Ife Adebara in her 2025 paper for the Centre for International Governance Innovation, the term language data flaring is the wasting away of language resources, particularly African language data - akin to gas flaring.
With gas flaring, it began because the gas couldn’t be poured into barrels, couldn’t be transported to distant cities, and had no immediate market. Flaring seemed like the cheapest and safest way to get rid of it to keep lucrative oil flowing.
With language data flaring, it is the “undercollection, poor storage, and limited use of Africa’s language data in AI systems”,locking away access to massive, immediate markets.
While the gas flaring problem is a loss that can be tapped and stopped, the language data flaring problem is more like a weed that deepens its roots as it grows. The longer policies struggle to catch up, the tougher it becomes to upend.
Does colonial legacy have anything to do with language low-resource status?
A language’s resource status is rarely created by one factor. There are always multiple contributors, and colonial legacies across Africa created a head start.
Post-colonial independence, former colonial languages continued to be prioritized, making them the sole or primary languages used in administration, education, media, governance and global integration. This preserved a colonial practice and created lasting hierarchies,relegating many African languages to oral use primarily and limiting their written development.
Why some African languages are better-resourced than others
Of Africa’s 2,000+ languages with varied but similar histories, languages like Amharic, Swahili, and Afrikaans are considered better-resourced compared to the over two thousand other languages. And this did not happen by chance.
Afrikaans, an African language of West Germanic origin, developed primarily from the Dutch dialects of settler colonialists in South Africa. It was made an official language with full state support during apartheid and after. The government invested in the creation of standardized writing systems, literature, media, and use in education. The language spread to other parts of Africa such as Namibia and Botswana, increasing its speaker populations.
Swahili, before contact with colonial influences, was already a trade language widely spoken along the East African Coast. It was supported by colonial groups-the Germans and the British-for administrative and missionary interests. The Germans invested in replacing the Arabic script the language was written in at the time with the Latin script, and the British standardized the language across the territories.
Amharic, the official language of Ethiopia, is written in the Ge’ez or Ethiopic script. It has a long-standing literary tradition and has been the language of the court, administration, and the dominant Amhara ethnic group for centuries. Even though Ethiopia is home to over 80 languages, Amharic was promoted in government, media, and education.
The one thing Amharic does not share with Afrikaans and Swahili
The one thing Amharic does not share with Afrikaans and Swahili is colonial influence. Other than that, all three languages receive state backing and promotion. Yet, Amharic is relatively less-resourced compared to Afrikaans and Swahili.
Afrikaans has a speaker population of about 20 million. Swahili has a speaker population of over 200 million. Amharic has a speaker population of about 35 million. Yet, Afrikaans is better resourced compared to Amharic and Swahili. For many African languages, the speaker population does not have much effect on language resource status.
Beyond historical colonial pressures, there is a correlation between the resource state of an African language and maintained colonial links.
Amharic is written in the Ge’ez script, organically developed over centuries. Both Afrikaans and Swahili use the Latin script, which is currently the most widely used writing system, and the script most extensively represented in the current AI training data. Translating Amharic correctly requires custom tokenizers, which adds complexity to the process.
Afrikaans also shares significant linguistic structures with English and Dutch, making translation easier for AI systems already trained in these high-resource languages. Both Swahili and Amharic have distinct linguistic structures.
Additionally, Amharic is largely confined to Ethiopia, and its usage faces competition from English, a primary medium of instruction in its secondary schools and universities.
The use of English and other former colonial languages internationally appears to be a factor when it comes to promoting indigenous languages. For a low-income country like Ethiopia, trade and contact beyond its borders are largely tethered to the use of English.
While colonial legacies contribute to African language resource status, state backing (colonial/post-colonial), promotion, investment, and economic incentives remain core drivers.
Amidst neglect, progress is happening—but is it fast enough?
Forty-nine African governments have endorsed the AI declaration to build governance frameworks that are inclusive, ethical, and tailored to African realities. This includes data sovereignty and use of AI to serve development priorities in health, agriculture, education, and governance. Also, about 22 countries have published their AI strategies, and 21 are currently in the drafting stage. But ending language data flaring requires so much more.
In North Africa, Egypt has taken the lead with the launch of Karnak, a national LLM focused on Arabic dialects, while Nigeria launched N-ATLAS to support major local languages. But true digital sovereignty requires physical infrastructure. Setting a new benchmark for the African Union’s vision, Ghana’s April 2026 National AI Strategy committed $250 million to establish a sovereign computing center. This allows local data to be trained domestically rather than rented abroad. Furthermore, the strategy explicitly mandates curating one trillion tokens of Ghanaian-language data by 2030 - a direct, infrastructural countermeasure to flaring.
The AU’s strategy calls for the building of high-quality datasets, inclusion of African languages, and efforts to reduce bias from use of non-representative data. The national strategies of many nations prioritize AI’s use for immediate economic growth and sector applications much more than the inclusion of the different languages. Egypt’s Karnak LLM focuses on Arabic, and Nigeria’s N-ATLAS covers only three major languages.
The targeted initiatives and partnerships front has shown more on-the-ground results
WAXAL released a large open dataset with 1,846 hours of transcribed speech, 565 hours for text-to-speech, and a total of 11,000 hours of raw audio. The dataset spans across 27 Sub-Saharan African languages spoken by over 100 million people. LINGUA Africa, an initiative launched by Masakhane, is supporting the development of datasets in 50 languages. African Next Voices released 18,000+ hours of speech-to-text across 24 languages.
There are a host of other efforts, but they all remain driven by corporate sponsorships, foundations, startups, researchers, and community efforts. Sustainable, scalable interventions for 2,000+ languages across all 54 countries require a synergy between the governments and the communities that is yet to take shape. Tangible progress across dozens of languages remains at an early stage with limited real-world deployment at a population scale or integration into critical sectors. This leaves a big space amidst the fast-growing need for language utilization across sectors.
The ongoing 2022 Meta lawsuit filed by Abrham Meareg, the son of the Ethiopian academic Meareg Amare who was assassinated in his home following death threats posted on Facebook, remains a clear reminder of some of the risks making this urgent. AI moderation tools used by platforms to moderate hateful speech miss violence-inciting posts when written in low-resource languages.
Delayed integration has left room for security risks, alongside development bumps, and economic losses worth billions of dollars across sectors such as education, agriculture, finance, and health.
Language data flaring remains a Day 2 Problem—Lessons from Day 1
The gas flaring practice was inherited globally in the commercial oil industry. However, countries like Norway and the UK have managed the loss and hazard from flaring with stringent fines and infrastructure development.
While countries like the U.S. and Azerbaijan still face flaring, they have adopted approaches to fit their realities and reduce flaring over the years. Nigeria’s shift from punitive regulations only to an inclusion of avenues for investors has also reduced flaring in recent years compared to the early 2000s, even though flaring remains a challenge.
The experience from managing gas flaring does not need to be learnt anew. The approaches to managing language data flaring must stay agile, progressive, and tailored to the evolving realities on the ground.
What we can infer about the future from today’s policies
Across the continent, governments have shown continued efforts to invest in sovereignty mechanisms. However, frameworks remain mostly aspirational with less detailed paths to implementation. Without accelerated execution, the future is likely slow, fragmented progress with a few languages while most languages continue to waste away.
For sovereignty policies to catch up with the reality on the ground, where less than 1% of Africa’s language data is used in training AI systems, governments must shift to building an environment that supports investments and collaborations.
This includes sustained commitment to build and maintain datasets, train and deploy models, create usable tools, and support infrastructure and skills development.
There should be efforts to prioritize the integration of language data into national digital public infrastructure and create national repositories with clear access. However, standardized guidelines without enforcement are merely suggestions. Upgrading existing regulatory bodies such as Ghana’s move to transform its Data Protection Commission into a powerful Responsible AI Authority demonstrates how local ministries must actively interpret the 2026 AU guidelines to secure actionable legal enforcement over data ownership, community consent, and algorithmic transparency
South Africa has made tangible progress on the infrastructure front with SADiLaR (South African Centre for Digital Language Resources). This national research infrastructure, focused on the country’s official languages, partners with academic institutions to build text, speech, and multimodal resources. These resources have supported the creation of MzansiLM, an AI model built by University of Cape Town researchers.
This stride shows what sustained investment can make possible on the infrastructure building, integration, and skills development fronts.
Should African governments take a leaf from pre-colonial imperialists?
Many sovereignty policies risk flaring most languages by default, treating linguistic diversity as a cost to be managed, rather than a resource to be harnessed. As with gas flaring, the goal should remain keeping flares as low as possible.
This doesn’t change the fact that digitizing languages for AI is expensive. High-quality data collection, standardization, annotation, and model training require significant investment in infrastructure, management, time, and talent. This can make prioritizing a few languages seem like the feasible, sensible approach—very much like how gas flaring made sense in its early years.
But in a time when contributing language resources to AI systems holds the opportunity for substantial returns in trade, productivity, and innovation, patterns become useful heuristics.
Salikoko Mufwene, a Congolese-American linguist, shared possible explanations for Africa’s language diversity. As quoted in The Christian Science Monitor, “Traditional African kingdoms were not as assimilationist as the European empires…say the kings relied on interpreters to translate to them what was coming from territories that they ruled but where people spoke different languages, there is no particular reason why we should be surprised that there are so many languages spoken in Africa.”
Pre-colonial African traders and polities offer an interesting parallel. They moved and traded across the Sahel, coasts, Nile, and Sahara, engaging with linguistically diverse communities, often relying on interpreters, multilingual locals, and evolving trade languages. This facilitated economic exchange despite the diversity.
Today, the AI layer offers a modern equivalent of those interpreters with the capacity to support translation, connection, and communication across the continent’s vast array of languages. Fully harnessed, local language AI could largely reduce barriers in rural agriculture (through advisory services, pest detection, and market information), cross-border trade, public health, and public service delivery. This also offers the opportunity to preserve indigenous knowledge across traditional medicine, ecological practices, and oral histories opening up access and monetization avenues in the digital economy.
Pre-colonial examples show that diversity can be managed, harnessed, and prioritized in tiers alongside the focus on major languages. Given resource constraints, speaker-population economics, and technical limits, community involvement and focus on high-value use cases offer a pragmatic starting point.
The depletionist or conservationist path: What road is being travelled?
Africa is currently on a mostly depletionist path, but not too far past the crossroads. While existing conservationist pockets are making progress, they remain insufficient. And language data flaring continues at scale, in spite of some promising efforts trying to curtail it.
Without an uptick in sustained political will, coordinated funding, and implementation, the direction may remain largely unchanged.
Even though sovereignty policies hold the potential to change the trajectory towards conservation, the future of Africa’s language resources remains entombed in the space between vision and execution.
Key Takeaways
- Language Data is Being Extracted and "Flared": Foreign AI models are extracting and wasting African language data without equitable compensation, creating severe economic barriers for millions.
- True Sovereignty Requires Physical Infrastructure and Legal Teeth: Aspirational policies aren't enough; stopping data flaring requires investing in local computing centers and enforcing strict data ownership laws.
- Linguistic Diversity is an Economic Asset, Not a Technical Burden: Governments must treat linguistic heritage as critical infrastructure, using local language AI to drive trade and digital power rather than viewing it as a cost.