Google Working with Wikipedia to Translate 'Smaller Languages'

Google said on 14 July that it is working with Wikipedia contributors, translators and "Wikipedians" across India, the Middle East and Africa to translate more than 16 million words for Wikipedia into Arabic, Gujarati, Hindi, Kannada, Swahili, Tamil and Telugu.

Wikipedia has the common languages covered. English, German and French run to millions of articles between them, maintained by large volunteer communities. The languages named above have far fewer, and many of the entries that do exist are stubs: a sentence, a date, a link, nothing a reader can actually use.

The motive is not mysterious. Google and the Wikimedia Foundation have a shared interest in organising information and making it consumable for web users everywhere. Google needs indexable content in the languages of its fastest-growing markets. Wikipedia needs articles. The method, though, is more interesting than the motive.

How the articles were chosen

Google Product Manager Michael Galvez explained the process for Hindi. "First, we used Google search data to determine the most popular English Wikipedia articles read in India," he said. "Using Google Trends, we found the articles that were consistently read over time, and not just temporarily popular. Finally we used Translator Toolkit to translate articles that either did not exist or were placeholder articles or stubs in Hindi Wikipedia."

The result: in three months, a combination of human and machine translation tools produced 600,000 words from more than 100 English Wikipedia articles, growing Hindi Wikipedia by almost 20 percent. Google then washed, rinsed and repeated for the other languages, bringing the total to 16 million words.

That selection step deserves attention. Using Google Trends to filter for sustained rather than spiky interest is a sensible way to avoid spending translation budget on a footballer who was famous for a fortnight. It also means an encyclopaedia's coverage in Hindi is being shaped by what Indian users searched for in English, which is a subtly different thing from what Hindi readers want to read.

Human plus machine, not machine alone

The phrase "a combination of human and machine translation tools" is doing a lot of work. Translator Toolkit produced a first draft, and human translators and Wikipedians cleaned it up. This is post editing, and for a factual encyclopaedia article it is close to the ideal use case: the source text is plain, the terminology is stable, and the goal is comprehension rather than style.

Hindi translation at this scale would be prohibitive if done from scratch. Draft-then-edit made it affordable. The same logic applied to tamil translation and swahili translation, where the pool of available professional translators is much thinner and the cost per word climbs accordingly.

The catch is that the quality of the draft depends on how much training data the engine has seen, and low-resource languages are low-resource precisely because that data does not exist in volume. A weak draft can be slower to fix than an empty page, a complaint that professional translators have been making for years and that the industry has argued about ever since. The case that machine translation cannot substitute for a human is at its strongest exactly where the data is thinnest.

Why these seven languages

The list is not random. Arabic, Gujarati, Hindi, Kannada, Swahili, Tamil and Telugu together cover well over a billion people, and almost none of that population is well served by English Wikipedia. They are also languages where Google had a commercial reason to act: India and East Africa were among its fastest-growing user bases, and a search engine is only useful if there is something worth finding at the other end of the results page.

The technical term for what these languages have in common is a shortage of parallel text. A translation engine learns by reading the same document in two languages, millions of times over. For French and English, the European Union alone supplies decades of it. For Kannada and English, there is far less, and what exists is often narrow: religious texts, government notices, news wire copy. An engine trained on that diet handles a legal notice competently and an article on quantum mechanics badly.

Feeding translated Wikipedia articles back into the public domain quietly fixes part of that problem. Every reviewed article becomes training material for the next generation of engines, which is the least discussed and possibly most important consequence of the whole project.

What the volunteers thought

Wikipedia's communities are not passive recipients of donated content. Editors on smaller-language Wikipedias have historically been wary of bulk imports, because a flood of machine-drafted articles creates a maintenance burden that falls on a handful of unpaid people. A stub is a small problem. Two thousand fluent-sounding articles with subtle errors is a large one, and the debate flares up regularly on forums such as r/wikipedia.

Some language communities have gone further and restricted or banned unreviewed machine-translated contributions outright, on the grounds that an encyclopaedia nobody trusts is worse than one that is merely small.

What it changes

There was a time when information barriers were inherent and assumed, thanks to language gaps. That time is closing. A Swahili speaker with a phone can now reach a body of reference material that did not exist a decade ago, and the marginal cost of adding the next language keeps falling.

There is something exciting and a little unnerving about the universalisation of Google and Wikipedia. Two organisations now decide, in effect, what a large share of the planet can look up, and the selection algorithm is a search trends report. Google still has a long way to go, because machine translation is an imprecise practice and a tough nut to crack. The 16 million words are a start. Whether the articles are read, corrected and kept alive by the communities they were written for is the part no engine can automate.