Translating the Internet
From its earliest days the Internet carried a promise: information moving at the speed of light, geography made irrelevant, anyone who shared your interests within reach. One old obstacle stood in the way of that, and it has not moved. It does not matter what you have access to if you cannot read it.
For the first couple of decades the web had a simple solution, and an unsustainable one. Most people online spoke English, or wrote in it anyway. That arrangement is now falling apart, which is good news wearing the costume of a problem.
The moment the web stopped writing in English
Ethan Zuckerman, founder of the multilingual blog network Global Voices, watched the shift begin. In 2004 he had dinner in Amman with a couple of dozen bloggers who spent the evening chatting in Arabic. "But almost all of them were blogging in English at that point," Zuckerman says.
Years later he went back. "Out of that group of people that I had dinner with, a lot of those people blog in Arabic now," he says. One of them explained why. "When we were trying this in 2004 there were very few Arabic speakers online, and we just couldn't write for that audience. But now our friends, our peers, our neighbors are all online. That's who we want to reach."
The numbers back the anecdote. According to Internet World Stats, the number of Arabic-speaking internet users rose by more than 2,000 percent over a decade. Chinese is on course to overtake English as the most-used language on the web. Dozens of other languages are growing fast.
Nobody writes for an audience that is not there. Once the audience arrives, people write for it in the language they think in. That is not a retreat from the global web. It is what a global web actually looks like.
The fracture risk
Which raises the awkward question. If everyone writes in their own language, does the internet split into parallel webs that never touch? A billion new users is a gain for them and a loss for the rest of us if none of us can read what they publish. The language barrier does not disappear when more people come online. It moves inside the network.
Google's answer: read the whole web
Much of the hope rests on machine translation, a technology with decades of mediocre history behind it and a sharp recent improvement. Google's approach was to stop teaching the computer grammar and start showing it text.
"What we do is use hundreds of billions of words that Google infrastructure has access to," says Michael Galvez, a project manager at Google Translate. The company's machines crawl the web, ingest the text, analyse it and learn how people actually write, then combine that with high-quality translated transcripts. Run a Spanish newspaper article through it and the English that comes back is genuinely usable.
Galvez is careful about what usable means. "Google Translate is good at helping you get what is called a gist, or essentially the essence of what the other person is communicating," he says.
The gist is a real product and it is not nothing. It is also not enough. Much of what is worth reading online is written with nuance, humour, register and rhythm, and software strips those out first. Some language pairs work far better than others, and even a good output is never quite right.
Meedan: keep both languages on screen
So some projects went back to humans. Meedan.net is one. "The idea is a Wikipedia-style approach to translation," says founder Ed Bice. Meedan mixes human and machine translation to present articles, blog posts and comments about the Middle East, aiming to connect the Arabic and English-speaking worlds.
Its most interesting decision is not about accuracy. It is about layout. Google Translate erases the original: you see the page in your language and the source is gone. Meedan puts English and Arabic side by side. Once you can see comment threads bouncing between the two, the translation stops being a verdict and becomes an invitation. A reader who speaks both can check it. A reader who speaks one can see that another language was there.
That design choice speaks to something the accuracy debate misses. Presentation of translated text matters as much as the text. Hiding the source implies the translation is the truth. Showing it admits that a person made choices, and lets someone else disagree, which is roughly how communities like r/languagelearning approach the same material.
No universal translator yet
People who study this expect both machine and human translation projects to keep improving quickly. Almost nobody will predict when, or whether, a Star Trek style universal translator arrives.
What is already clear is the direction. The web is moving away from English and it is not coming back. Global Voices and Meedan both started from the same premise: a translated internet is not a technical problem with a technical answer, but a mixed system where software does the volume and people do the meaning. Anyone reading the web over the next decade will be leaning on both, and the honest ones will keep saying so.