Computer programs from the University of Leeds could change how languages are taught

Language teaching has a supply problem that nobody in a classroom can solve. Textbooks take years to write and years more to revise, and by the time a vocabulary list reaches a student it describes a language that has moved on. Modern language experts at the University of Leeds have been building computer programs aimed squarely at that lag, and the tools they produce could change what translators and learners work with day to day.

Dr Serge Sharoff of the Centre for Translation Studies is running three research projects at once: Kelly, TTC and Accurat. The European Commission has awarded him close to £700,000 to fund the work. That is a serious sum for research that never leaves the level of words, and the reason is straightforward. Words are where every other language technology begins.

What TTC actually does

TTC stands for Terminology Extraction, Translation Tools and Comparable Corpora, a name that hides a practical goal. It aims to produce current, usable term lists drawn from original texts in French, German, Spanish, Chinese and Russian, targeting fields that move faster than any dictionary can follow. Wind energy is one. Mobile phones are another.

Consider what a translator faced with a wind turbine maintenance manual is up against. The vocabulary of pitch systems, yaw drives and gearbox condition monitoring did not exist in any glossary a decade before the industry scaled. A translator either invents a term, guesses at the industry's preference, or spends unpaid hours reading Spanish engineering papers to find out what practitioners actually say. Multiply that across every emerging sector and the cost is enormous.

TTC's approach is to mine that answer automatically. The method draws on a variety of techniques applied to large volumes of original text in each language. The resulting term banks feed directly into the tools translators already use, including machine translation systems and computer-assisted translation software, and into multilingual content management platforms that need consistent terminology across thousands of documents.

Comparable corpora, and why they beat parallel texts

The technically interesting part sits in the word "comparable". Traditional statistical translation systems were trained on parallel corpora: the same document in two languages, aligned sentence by sentence. Those are rare, expensive and heavily skewed towards the material that institutions happen to translate, which in practice means European Union legislation and Canadian parliamentary proceedings.

Comparable corpora loosen the requirement. Instead of the same text twice, you gather texts in two languages about the same subject, written independently. Spanish wind energy trade press on one side, German wind energy trade press on the other. Neither is a translation of the other. Statistical terminology extraction then identifies which terms occupy equivalent positions in each language's usage patterns, and pairs them.

The advantage is scale. Comparable text exists for almost any subject in almost any language, because people write about their own industries in their own languages all the time. That is the raw material corpus linguistics has been refining methods for since the 1960s, now pointed at a commercial problem.

The Kelly angle: teaching what people actually say

Kelly addresses the classroom side. Vocabulary lists in language courses have long been assembled by intuition, by tradition, or by copying the previous textbook. Corpus-derived frequency lists replace that with evidence. If you rank words by how often they genuinely appear in contemporary usage, you can tell a beginner precisely which two thousand words will carry them through most of what they will read and hear, and in what order to learn them.

The result is unglamorous and highly effective. It also exposes how much classroom vocabulary is decorative. Generations of students have learned words for the parts of a farmyard while lacking the vocabulary to open a bank account.

Why translators should care

Sharoff's projects sit upstream of the tools most working linguists touch. Term banks are what populate the glossary pane in a translation environment, and glossary quality determines whether the software helps or hinders. Teams that have wired machine output into their CAT tools properly already know that the terminology layer, not the engine, is where consistency is won or lost.

The practical consequences look like this:

  • Faster onboarding for new domains. An extracted term list gives a translator a working vocabulary in days rather than months.
  • Consistency across large teams. Ten linguists on one manual will produce ten synonyms unless a shared term bank stops them.
  • Better machine output. Feeding verified terminology into an engine constrains it where it is weakest, on domain-specific nouns.
  • Cheaper maintenance. Re-running extraction on fresh text updates a glossary without a human rereading the field.

Public money, public tools

The European Commission funding is not charity. Europe runs on multilingual documentation, and the cost of producing it is measured in billions. Research that shaves time off terminology work pays for itself across the institutions that commissioned it, then spills out to everyone else, because the outputs of projects like these tend to be published rather than locked away.

Whether the tools change how languages are taught is a slower question. Curricula move at their own pace, and evidence about word frequency has to fight tradition. But the direction is set. Practising translators trade notes on exactly this on r/TranslationStudies, where the perennial complaint is not that machine translation is bad, but that it is confidently wrong about the one word that mattered. Fixing the words fixes a good deal else. The Centre for Translation Studies at Leeds is betting on that, and the Commission is paying for the bet.