# Nobody asked the language industry

Date: 2026-09-23
Author: Richard Brooks
Canonical: https://richard-brooks.com/nobody-asked-the-language-industry/

On Monday in New York, during [the week the UN takes over the city](https://www.un.org/en/ga/), [sixty organisations put their names to a five-year goal](https://www.gatesfoundation.org/ideas/media-center/press-releases/2026/09/ai-language-partnership). The idea is that by 2031, the 3.4(ish) billion people who speak languages that today's AI handles badly should be able to use it in their own language and voice. The [Gates Foundation](https://www.gatesfoundation.org/) announced it. Bill Gates was there for [Goalkeepers](https://goalkeepers.gatesfoundation.org/about-event/), his foundation's annual gathering, at [Jazz at Lincoln Center](https://www.jazz.org/).

![The United Nations Secretariat Building on Manhattan's East River, with the Chrysler Building and the Midtown skyline behind it](/images/un-headquarters-new-york.webp)

I read the list of signatories.

The frontier labs are all there: [Anthropic](https://www.anthropic.com/), [Google](https://ai.google/), [Microsoft](https://www.microsoft.com/en-us/ai), [Amazon](https://www.aboutamazon.com/), [Mistral](https://mistral.ai/), [NVIDIA](https://www.nvidia.com/), the [OpenAI Foundation](https://openaifoundation.org/). So is [ElevenLabs](https://elevenlabs.io/) (they do voices), and [Zoom](https://www.zoom.com/) (surprised). Then the people who have been doing the actual work for years, mostly on grants and goodwill: [Masakhane](https://www.masakhane.io/), the grassroots African NLP network; [AI4Bharat](https://ai4bharat.iitm.ac.in/) and [BHASHINI](https://bhashini.gov.in/) in India; [Lelapa AI](https://lelapa.ai/) in South Africa; [Digital Umuganda](https://digitalumuganda.com/) in Rwanda; [Karya](https://www.karya.in/), which pays rural Indians to record their own languages. Then the money and the state: the [World Bank](https://www.worldbank.org/), [UNICEF](https://www.unicef.org/), the [UK Foreign Office](https://www.gov.uk/government/organisations/foreign-commonwealth-development-office), [Senegal's digital ministry](https://www.mctn.sn/), Canada's [IDRC](https://idrc-crdi.ca/en). Oh, and [an archbishop](https://en.wikipedia.org/wiki/Vincenzo_Paglia).

Here is who is not on it. No translation company. No interpreting firm. No localisation provider of any size. No industry association. The language industry, with its relatively tiny valuation of $50bn to $100bn, has been putting words into other languages for money since before any of these labs existed. It is not in the room.

I don't think that was a snub. I think nobody thought to ask. Or who to ask. That is an interesting problem.

## Mind the gap

There are around [7,000 languages](https://www.ethnologue.com/insights/how-many-languages/). A small fraction of them have enough written material online to train a good model. The Gates Foundation says [more than 90 percent of the data](https://www.gatesfoundation.org/ideas/media-center/press-releases/2026/09/goalkeepers-report-equitable-ai) behind the early large language models was English. E.M. Lewis-Jong, who runs the [Mozilla Data Collective](https://community.mozilladatacollective.com/about/) (another signatory), [put it to AP](https://abcnews.com/Technology/wireStory/gates-foundation-launches-coalition-build-representative-language-data-136622214) as a question: "Why would you think that you were going to get a culturally diverse and representative system out of something that was predominantly trained on Reddit?"

The [foundation's own report](https://goalkeepers.gatesfoundation.org/wp-content/uploads/2026/09/2026_Goalkeepers_Report_EN.pdf) has the example that should end the argument. A woman in labour in Malawi says, in Chichewa, that her water has broken. Translated word for word into English, she has "thrown away water". That is not a cute error. It is the difference between someone getting to hospital and someone not.

So the goal is right. What I want to look at is what the announcement actually commits to, and the potential commercial interest.

## The target

This is sixty signatories. The press release names no sum of money for the coalition. It says each organisation will contribute "their individual resources and expertise", and that "the coalition's detailed structure, governance, and workstreams will be developed collaboratively over the coming year".

The money came the week before. On 14 September the Gates Foundation [committed $1 billion to AI over the next two years](https://qz.com/gates-foundation-billion-ai-global-health-education-agriculture-091526): 40 percent to education, 40 to health, 10 to agriculture, and 10 to what it calls the digital foundation, which includes [datasets in underrepresented languages](https://www.semafor.com/article/09/21/2026/gates-foundation-commits-1b-to-local-language-ai-development-in-africa). So roughly $100 million over two years, shared with other foundation work, across a problem the size of 3.4 billion people. That is serious money for a grant programme. It is not much for a market.

What Monday produced is a shared target, which in its way is more powerful than money. It has four workstreams:

1. **The open language layer.** Shared data infrastructure "every builder can draw on using open licenses".
2. **Tracking progress honestly.** Assessments and benchmarks that measure real gains against the goal.
3. **Working tools.** Models and applications "usable by any AI builder, not just those with the most resources".
4. **Reaching people safely.** Privacy, consent and data sovereignty.

Read those as a buyer would, not as a campaigner would. The first one sets a price (free). The second one sets a standard (theirs). Those are the two things that decide who wins a contract.

## Underrepresented is not unserved

The press release calls these languages underrepresented, and in a model's training data they are. In the real world plenty of them have been served for decades, by people, but at a price.

When a hospital needs a Tigrinya interpreter at three in the morning, somebody finds one. When a court needs a document in a language of the Sahel, a firm somewhere has a freelancer on its books. When an aid agency runs a vaccination campaign, someone translates the leaflet into the languages the vaccinators will actually meet. Most of it is done by small firms and individual linguists on thin margins, and in the rarer languages it is expensive, because the supply is a handful of people and the demand is lumpy.

I have been treasurer of both [ELIA](https://elia-association.org/) and [ALC](https://www.alcus.org/), the European and American associations of language companies. I have sat through a lot of budget meetings about exactly this end of the business. The rare-language work was never the profitable part. It was the part you did because the client needed it and nobody else could.

So the work the coalition wants to do was already being done. It was expensive, fragmented, and invisible to anyone who measures a language by how much of it exists on the internet.

There is a good example from Canada this summer. [Winnipeg police started trialling](https://www.cbc.ca/news/canada/manitoba/artificial-intelligence-translation-police-cameras-9.7320248) Axon's body-camera translation in August. [Axon's list](https://www.axon.com/resources/axon-assistant-faqs) runs to more than fifty languages. It has Welsh and it has Māori. It has no Cree and no Ojibwe. Fifty languages, and not the ones on Winnipeg's doorstep. That is not malice. It is how a market works: coverage follows volume. Gates said as much on the 14th: "Left to the market alone, the most capable tools will be built first for the people and institutions most able to pay for them."

He is right. But the language industry is that market. It is the bit of it that already goes where the volume isn't, when someone pays.

## Where the industry's money actually is

Traditional translation is more or less flat. [Nimdzi](https://www.nimdzi.com/nimdzi-100-2026/) has the whole industry growing at under 1 percent a year to 2030. [RWS](https://www.investegate.co.uk/announcement/rns/rws-holdings--rws/half-year-financial-report/9612324), the biggest UK-listed company in the sector, reported in June that its core language division shrank 4.6 percent. Its Generate division, home of the TrainAI data business, grew 52 percent to £99.4 million, 28 percent of the group. Nimdzi found that data services (collection, curation, evaluation) were the biggest source of revenue growth and margin protection across the top hundred providers last year. It puts the data arms of TransPerfect, RWS, Lionbridge and Welo Global in the same market as [Scale AI](https://scale.com/).

Lionbridge [sold its AI data unit to TELUS](https://slator.com/lionbridge-sells-ai-division-to-telus-for-usd-935-million/) for $935 million back in 2020.

The fastest-growing thing the language industry sells is language data for AI: recordings, transcriptions, translations, annotations, human ratings. In the long tail of languages that is the premium end. Few people can do the work, the labs need it, and they have almost unlimited money to spend.

The coalition's first workstream is an open, shared, openly licensed language layer, funded by philanthropy and co-signed by the labs. The biggest buyers of language data have joined a coalition to make a large part of it free. That is a sensible move for the labs, and it may well be good for the 3.4 billion people who speak these languages. If you run a language company and your growth plan includes low-resource data collection, Monday's news changes your price/services list. Free data still has to be collected with consent, checked by people who speak the language, and made to work in a hospital, a court or a classroom. That work needs to be paid for.

## Open to whom

The other thing "open" does is move value to whoever can do the most with free inputs. That means the people with the compute.

In 2023 Te Hiku Media, the Māori broadcaster that built its own te reo Māori speech recognition, published a blog post called ["OpenAI's Whisper is another case study in Colonisation"](https://blog.papareo.nz/whisper-is-another-case-study-in-colonisation/). Whisper had been trained on 1,381 hours of te reo Māori, apparently scraped from the web. Their question was simple: "who gave them the right to create a derived work from that data and then open source the derivation?" Te Hiku keeps its own data under a [Kaitiakitanga licence](https://tehiku.nz/te-hiku-tech/te-hiku-dev-korero/25141/data-sovereignty-and-the-kaitiakitanga-license), which only allows use that benefits Māori people. OpenAI's foundation is now on Monday's list. Te Hiku is not.

Karya, which did sign, takes a different route to the same destination. It [pays its workers around $5 an hour](https://time.com/6297403/the-workers-behind-ai-rarely-see-its-rewards-this-indian-startup-wants-to-fix-that/), about twenty times India's minimum wage, and says that when a dataset is resold, the workers get the proceeds. Its founder, Manu Chopra, told Time: "you can't solve a market failure in the market."

Both are on the right side of the fourth workstream, the one about "privacy, consent, and data sovereignty". Both sit awkwardly with the first one, which is about open licences. The coalition has a year to work out how those two live together. The answer decides whether the open layer becomes a commons for the communities named in the release or a cheap training run for the firms that already have the GPUs.

## The benchmark will end up in the tender

The second workstream worries me in practice.

Benchmarks are a good thing. The industry has always measured itself with delivery rates, complaints and the account manager's reassurance. A shared measure of whether a model works in Yoruba beats a vendor's word for it.

But I teach key account management, and I know what happens to a good measure once procurement finds it. It becomes a box. Within a couple of years a health ministry tender, or a World Bank-funded programme, or a UNICEF brief, will ask whether your solution meets the coalition's benchmark for the languages in scope. Many of those buyers are on the signatory list. The firms that can answer with test results will win the work. ISO certificates and a list of linguists won't be enough.

And averages hide things. [Hugging Face added Hindi and Indian English](https://huggingface.co/blog/open-asr-leaderboard-global-south) to its Open ASR Leaderboard in August, with nearly 4,900 speakers tagged by region. Two systems that look the same on the leaderboard turned out to differ almost fourfold in how much their accuracy depends on where the speaker comes from. One score per language tells you very little about whether a system works for the person in front of you.

Someone has to do that measuring properly. Native-speaker evaluators, recruited by region and dialect, briefed and paid, with their judgements checked. That is a job, and the language industry already knows how to do it.

## What I would do if I ran a language company

**Map the list like a key account.** Sixty organisations, a five-year goal and a year of governance still to design. Sort them. The labs will compete with you for some work and buy from you for the rest. The funders and agencies, UNICEF, the World Bank, the FCDO, [CHAI](https://www.clintonhealthaccess.org/), will write the tenders. The platforms that deliver the services, such as [Digital Green](https://digitalgreen.org/) for farmers and [Viamo](https://viamo.io/) for information by mobile phone, are the ones that need the languages to work. In each organisation find three people: whoever owns the language question, whoever holds the budget, and whoever will sign off the benchmark. Get to them this year, while the structure is still being designed.

**Put agents on watch.** Build a small team of AI agents to follow the coalition's announcements, papers, job posts and social chatter, and brief your business development team every morning. My own news desks run on [Grok](https://grok.com/), and [Hermes Agent](https://hermes-agent.nousresearch.com/) is another option. The firm that hears about the first working group or pilot tender a week early gets the first meeting.

**Pitch against what each buyer measures.** UNICEF counts children vaccinated. A health platform counts referrals that reach a clinic. An agriculture service counts farmers who act on the advice. Use those numbers to show what your language work is worth to them in their language. If a campaign message in the mother tongue gets 30 percent more mothers to the clinic than the same message in the official language, those extra visits are your value, and your price should be set against them. Run the comparison on a pilot, measure it with the client, and put the result on the first page of your next proposal.

**Sell data that survives an audit.** Consent that stands up, provenance you can trace and a clear answer to "who benefits?" will be scarce, and the fourth workstream makes them a requirement. So will domain depth. A clinical glossary in Chichewa is worth more to a hospital than ten thousand hours of Chichewa radio.

**Partner with the community groups.** Masakhane, Karya and Digital Umuganda have the relationships and the trust. You have project management, quality systems and paperwork that a ministry's procurement team recognises. Together you can bid for work neither of you would win alone.

**Build the evaluation business now.** "Tracking progress honestly" needs native-speaker raters, recruited by region and dialect, briefed, paid and checked, in hundreds of languages. The best firms have been running quality management in rare languages for twenty years. Test the models in the languages you know best, publish the results, and be the firm the coalition calls when it needs raters.

**Sell the review.** Deployments in health, courts and public services will need a qualified human between the machine and the person it could hurt. The liability sits there, and so does the price.

**Get the associations in the room.** ELIA, ALC, [GALA](https://www.gala-global.org/) and the interpreters' bodies should ask to join. The coalition says its structure will be worked out "collaboratively over the coming year", and the profession that has served these languages for decades belongs in that conversation. I would send the letter this month, with an offer attached: introductions to members' linguists in the languages the coalition will find hardest to reach.

## Who does the work

The announcement is good news and I hope it works. A lot of the organisations on that list have earned their place the hard way.

I have been in this business long enough to know that when rich institutions discover a market, the people who were already serving it are usually the last to hear. These languages have been served by hand, one expensive language at a time, by people nobody at Goalkeepers has met.

The coalition has a year to decide how it will work. The language companies that turn up in that year, with evidence, partners and a clear offer, will help build what comes next and get paid for it.

Sources: Gates Foundation, [Global organizations announce five-year goal to help more than 3 billion people use AI in their own language and voice](https://www.gatesfoundation.org/ideas/media-center/press-releases/2026/09/ai-language-partnership), 21 September 2026. Gates Foundation, [Gates Foundation commits US$1 billion to help build and deliver equitable AI](https://www.gatesfoundation.org/ideas/media-center/press-releases/2026/09/goalkeepers-report-equitable-ai), 14 September 2026. Gates Foundation, [2026 Goalkeepers Report: AI, equity, and the choice we can't delay](https://goalkeepers.gatesfoundation.org/wp-content/uploads/2026/09/2026_Goalkeepers_Report_EN.pdf), September 2026. AP, [Gates Foundation launches coalition for more representative language data sets for AI](https://abcnews.com/Technology/wireStory/gates-foundation-launches-coalition-build-representative-language-data-136622214), 21 September 2026. Quartz, [Gates Foundation pledges $1 billion for AI in global health and education](https://qz.com/gates-foundation-billion-ai-global-health-education-agriculture-091526), September 2026. Semafor, [Gates Foundation commits $1B to global AI development](https://www.semafor.com/article/09/21/2026/gates-foundation-commits-1b-to-local-language-ai-development-in-africa), 21 September 2026. [Nimdzi 100, 2026](https://www.nimdzi.com/nimdzi-100-2026/). RWS, [half-year financial report](https://www.investegate.co.uk/announcement/rns/rws-holdings--rws/half-year-financial-report/9612324), 11 June 2026. Slator, [Lionbridge sells AI division to TELUS for USD 935 million](https://slator.com/lionbridge-sells-ai-division-to-telus-for-usd-935-million/), 2020. CBC News, [Winnipeg police adding AI translation tool to body-worn cameras](https://www.cbc.ca/news/canada/manitoba/artificial-intelligence-translation-police-cameras-9.7320248), August 2026. Axon, [Axon Assistant FAQs](https://www.axon.com/resources/axon-assistant-faqs). Te Hiku Media / Papa Reo, [OpenAI's Whisper is another case study in Colonisation](https://blog.papareo.nz/whisper-is-another-case-study-in-colonisation/), January 2023. Te Hiku Media, [Data sovereignty and the Kaitiakitanga licence](https://tehiku.nz/te-hiku-tech/te-hiku-dev-korero/25141/data-sovereignty-and-the-kaitiakitanga-license). Time, [The Indian startup making AI fairer, while helping the poor](https://time.com/6297403/the-workers-behind-ai-rarely-see-its-rewards-this-indian-startup-wants-to-fix-that/), 2023. Hugging Face, [The Open ASR Leaderboard adds its first Global South language](https://huggingface.co/blog/open-asr-leaderboard-global-south), 28 August 2026. Ethnologue, [How many languages are there in the world?](https://www.ethnologue.com/insights/how-many-languages/)