First Turkish-Romany dictionary to be completed soon
Necat Cetin does not look like a man rescuing a language. He is the principal of the Ozbey Primary School in the Torbali district of Izmir, on Turkey's Aegean coast, and he is also a historian. For almost two months he has been working four hours a day, on top of running a school, to compile something that has never existed before: the first Turkish-Romany dictionary.
The work is unglamorous, and it is done the only way this kind of work can be done. Cetin has been interviewing Roma residents of the Caybasi, Pamukyazi and Kuscuburun neighbourhoods of Torbali, one word at a time, with a team of ten Roma collaborators. He has identified the Romany equivalent of 4,500 Turkish words so far. He is currently researching words beginning with the letter Y. When the team finishes, the dictionary will contain more than 5,000 words.
Why nobody had done it before
The absence of a Turkish-Romany dictionary is not an oversight. It is a consequence of how the Romani language has lived for a thousand years. Romani descends from an Indo-Aryan ancestor carried out of northern India, and it has travelled continuously ever since, absorbing Greek, Slavic, Romanian, Hungarian and Turkish material along the route. There is no single Romani. There is a family of dialects, some barely mutually intelligible, each shaped by the majority language surrounding it.
Turkey's Roma communities speak varieties heavily inflected by Turkish. That makes a Turkish-Romany dictionary useful in a way a general Romani lexicon would not be, and it also makes it perishable. The vocabulary Cetin is recording belongs to people who are, in most cases, bilingual, and whose children are increasingly monolingual in Turkish.
What fieldwork actually looks like
Compiling a dictionary from living speakers rather than from texts is slow, because the source material walks around and has opinions. Speakers disagree. One neighbourhood uses a word another has never heard. Older informants remember terms the young have dropped. A lexicographer working from a written corpus can check a citation. A lexicographer working from a community must decide, repeatedly, what counts as the word.
Cetin's method is the standard one and it is punishing in its repetition: sit with speakers, work through a semantic field, record what comes back, cross-check with the next informant, then do it again. The team of ten Roma collaborators is not a courtesy. They are the only people who can tell him when a suggested equivalent is wrong, archaic, or a Turkish borrowing wearing a Romani coat.
The entries that never map cleanly
Anyone who has done this work knows the hardest entries are not the rare ones. They are the ordinary ones. Kinship terms carve up families differently across languages. Verbs of motion bundle direction and manner in ways that refuse to line up. Words for shame, obligation and social debt are notoriously resistant, which is why whole books of untranslatable words exist and keep selling.
A dictionary that pretends every Turkish word has one clean Romani equivalent would be a worse document than one that admits the mismatches. The 5,000-word target matters less than what the entries do with the awkward cases.
Why 5,000 words is a serious number
It sounds modest beside the hundreds of thousands of entries in a major reference work. It is not. For a language with no standard orthography, no state backing and almost no published literature, a 5,000-word bilingual lexicon crosses a threshold. It is enough to support teaching materials. It is enough to give written form to a speech community that has largely had none. It is enough, above all, to serve as evidence that the language is a system rather than a collection of household habits.
That evidential function matters more than it should. Communities whose languages lack documentation are routinely told, in effect, that they do not have a language but a way of speaking badly. A dictionary is a rebuttal you can put on a shelf and hand to a ministry.
The clock this is running against
Romani is not on the brink of extinction. It has millions of speakers across Europe. But individual varieties disappear quietly, and the mechanism is always the same. Parents make a rational calculation that their children will do better in the majority language, and one generation later the transmission chain has snapped. UNESCO classifies Romani as definitely endangered across much of its range for precisely this reason.
The pattern is well documented. In the final stage it is rarely persecution that kills a language. It is the accumulated decisions of people who love their children and want them to get jobs. Documentation projects such as the archive at the Endangered Languages Project exist to catch what those decisions leave behind.
What a schoolteacher can do that an institution cannot
The most interesting thing about this project is who is doing it. Not a university department with grant money and a research assistant. A primary school principal, four hours a day, with ten neighbours, in three neighbourhoods of a district most people outside Izmir could not place on a map.
That is not a lesser form of scholarship. Local documentation has advantages institutional projects rarely buy: trust, access, and the patience to keep coming back next week. Linguists arguing about method on r/linguistics tend to circle the same conclusion, which is that the best fieldwork usually involves someone the community already knows and has no reason to perform for.
Cetin will finish at Y, then Z, and then the thing will simply exist. Five thousand words that lived only in people's mouths will be on paper, which is where languages go when they need to survive the generation that stops speaking them.