Translation and Litigation
In litigation, knowing exactly what an email says is not a nice-to-have. It is the case. That gets difficult the moment the email is not in your language, and it gets expensive the moment there are 40,000 of them.
David Kessler, a partner at Drinker Biddle & Reath, remembers a matter involving a multinational company. Most of the discovery came back in English. Then the legal team noticed that a few of the company's employees had been emailing each other in an Eastern European language. It looked odd. They ran the messages through machine translation simply to get a sense of what was in them. Those messages turned out to be linchpins in the case.
The triage principle
Kessler took a specific lesson from that, and it is the most useful sentence anyone in this field has said about the technology.
"Machine translations are not very good at idioms, not very good in context, but they can be useful in terms of getting a sense of the document to let you decide if you want to spend more money," Kessler says.
That is the whole doctrine in one line. The engine is not there to produce a translation you can rely on. It is there to sort a mountain into two piles: probably irrelevant, and worth paying a human to read properly. Used that way, an imperfect output is not a defect. It is a filter, and the only quality it needs is enough fidelity to route the document correctly.
Where the money actually goes
A single dispute in a global economy can involve a variety of languages, and the cost of translating everything is what makes cross-border matters so brutal to budget. Human legal translation is priced per word, per language pair, and legal material sits at the expensive end because the translator has to understand what a clause is doing, not just what it says. Multiply that by a document population that has already survived collection and you are into six figures before anyone has read anything.
This is why corporations have leaned on translation technology and e-discovery platforms that support multiple languages. Handled well, the combination genuinely constrains budgets and wins cases. It lets a team cut a 40,000 document population down to the 800 that a qualified linguist and a qualified lawyer then work through line by line, which is where reliable legal translation services earn their fee.
The three ways it goes wrong
The drawbacks are real and they are all failures of process rather than of software.
- It fails to save time. Teams run everything through an engine, then discover the output is unreadable enough that reviewers slow down rather than speed up. Two passes are made where one would have done.
- It increases translation costs. Machine output gets sent for human post-editing on documents that were never relevant. Paying a linguist to clean up a machine translation of a lunch order is worse than not translating it at all.
- It misses documents in keyword search. This is the dangerous one. A team builds a search term list in English, applies it to machine translated text, and assumes coverage. But the engine may render a key term three different ways across a corpus, or use a synonym the search list does not contain, and the responsive document sits quietly outside the hit list.
That last failure mode is why search terms in a multilingual e-discovery exercise have to be built in the source language, by someone who speaks it, and tested against the actual corpus rather than against an English translation of it. Idiom, slang, internal shorthand and industry jargon all defeat a naive term list. Employees writing to each other do not use the vocabulary of a contract.
What careful practice looks like
The teams that get this right treat translation as a staged pipeline rather than a single purchase. Language identification runs first, because you cannot triage what you have not classified, and mixed-language corpora routinely contain documents nobody expected. Machine output then supports relevance decisions only, and nothing more. Anything that will be cited, quoted in a brief, put to a witness or filed gets a human translation, and the fact that a document was first surfaced by an engine is never a reason to skip that step.
Courts expect the same discipline. Under the Federal Rules of Civil Procedure, a party's discovery obligations do not shrink because the material is inconvenient to read, and the federal judiciary maintains its own certification regime for the interpreters who work in the courtroom precisely because untested language competence has been a source of reversible error. Practising litigators trade horror stories about this constantly in forums like r/Lawyertalk, usually variations on the same theme: the one document nobody bothered to translate properly.
Getting the languages counted first
One practical step separates the teams that control this from the teams that get surprised by it. Before any budget is set, run language identification across the whole collected population and produce a breakdown by document count and by custodian. It takes hours, it costs almost nothing, and it converts an unknown into a number. A matter that was scoped as English with a bit of German, and turns out to be 22 percent Mandarin, is a different matter with a different price, and the time to discover that is week one rather than the week before a production deadline.
The Kessler story is instructive because of how close it came to going the other way. A handful of messages in an unexpected language could easily have been dismissed as noise, or run through an engine, read as gibberish and set aside. Somebody thought it was odd and looked harder. Technology bought them the cheap first look. Judgement did the rest.