Thai Government Backs Wikipedia Localization to Widen Access to Knowledge
Thailand's government backed a project to expand the Thai edition of Wikipedia using statistical machine translation, feeding English articles through an engine and inviting volunteers to clean up what came out. The reaction split neatly along a line that has defined the machine translation debate ever since: those who saw a knowledge gap being closed, and those who saw a knowledge gap being filled with something that only looked like knowledge.
The problem the project was trying to solve
The scale gap between language editions of Wikipedia is enormous and self-reinforcing. English has millions of articles. Most languages have a fraction of that, and the ones they lack are frequently the technical, scientific and medical entries that a student most needs. A Thai speaker researching a physics topic in Thai often finds a stub, or nothing, and switches to English if they can.
That switch is the quiet inequality at the heart of the multilingual web. Access to knowledge ends up conditional on knowing a second language, which sorts opportunity by education before anyone opens a book.
Why machine translation looked like the answer
The economics are obvious. Human translation of a hundred thousand encyclopedia articles is a decade-long project requiring money nobody has. Machine translation produces a first draft in hours at almost no marginal cost, and the volunteer community can, in theory, correct it.
The theory has a weak point, and everyone who has run one of these projects has found it. Correcting bad machine output is often slower than translating from scratch, because the editor has to first work out what the source actually said, then diagnose what the machine got wrong, then rewrite. Volunteers who signed up to write encyclopedia articles rarely stay to do quality assurance on a machine.
The criticism was mostly about trust
The objections raised at the time were not really about grammar. They were about what happens when a reader cannot tell whether an article was written by a person who understood the subject or generated by a system that did not. An error in a Wikipedia article about a chemical compound or a medical dosage is not a stylistic problem.
Statistical systems of that era were fluent enough to be dangerous. They produced text that read plausibly and inverted meanings, dropped negations and hallucinated specifics. A fluent error is far harder to catch than an obviously broken one, and a volunteer skimming for typos will not catch it at all.
What the Wikimedia community learned
The movement's own approach evolved in response. Machine translation now sits inside the editing tools as an assistive layer rather than a bulk import channel, and several language communities have introduced limits on how much unedited machine output an article may contain before it is flagged or deleted. The Wikimedia Foundation has consistently treated translation as a way to help a human editor work faster, not as a substitute for having one.
That distinction has held up remarkably well as the technology has improved. Neural systems produce far better Thai than the statistical engines of the early 2010s, and the governance question is unchanged: who is accountable for the sentence, and does anyone who understands the topic read it before a student does?
Thai is a hard case
Thai puts particular pressure on machine systems. It has no spaces between words, so segmentation has to be inferred before translation can even begin. It has no inflection for tense or number, which means information that English states explicitly must be reconstructed from context. Register is encoded through pronouns and particles that carry social relationships English does not mark at all.
Providers of thai translation services deal with these problems daily in commercial work, and their answer is the same one the encyclopedia arrived at: use the machine for throughput and put a native speaker on anything that matters. Official documents, medical text and legal material are never signed off on machine output alone.
The measurement problem
One reason these debates go in circles is that the industry has never agreed on how to score the output. Automatic metrics compare a machine's sentence against a reference translation and reward overlap, which means a fluent sentence that reverses the meaning of the original can still post a respectable number. Human evaluation catches that, but it is slow and expensive, which is exactly what the project was trying to avoid in the first place.
The practical workaround used by most localisation teams is risk-based. Content is sorted by what happens if it is wrong. An article about a pop group can tolerate an awkward sentence. An article about drug interactions cannot, and it gets a human reviewer regardless of what the metrics say. Applied to an encyclopedia, that principle would have sent the machine at the long tail of geography and sport, and kept it away from medicine and law entirely.
The broader lesson for localisation
Every organisation that has attempted large-scale localisation has run the same experiment. The pattern that emerges is consistent:
- Machine translation is excellent for gist, search and internal documents where speed beats polish
- It is adequate for high-volume commercial content with light human editing
- It is unacceptable, unedited, for anything a person will rely on to make a decision
- The cost of post-editing bad output frequently exceeds the cost of translating well the first time
Translators debating this in professional forums such as r/TranslationStudies tend to be less hostile to the technology than outsiders assume. What they object to is not the tool. It is the assumption that the tool removes the need for someone who actually understands both languages.
Where it leaves the project
The instinct behind the Thai initiative was right. A knowledge commons that only works in English is not a commons. The mechanism was the part that needed rethinking, and the version that has endured across the encyclopedia's language editions keeps the machine in the loop and the human in charge.