AI Finally Speaks Yiddish

AI today speaks a lot of languages. English, obviously. Then French, Arabic, Hindi, Japanese — dozens more, sometimes a hundred or more in a single model.

But Yiddish? Ask one of the big AI chatbots a question in Yiddish and you'll often get something that looks like Yiddish the way a costume looks like clothing. The words are roughly right. The soul is missing. Idioms come out mangled, grammar bends in ways no bobe would allow, and the answers read like what they actually are: a translation of English thoughts wearing Yiddish spelling.

A small team of researchers just changed that. They released MameLoshnLM — what they believe is the first large AI language model built specifically for Yiddish — along with a paper explaining how they did it. "Mame-loshn," in case you haven't heard the word in a while, is Yiddish for "mother tongue." Literally: mother-language.

Here's the story of why it took until now — and why it matters far beyond Yiddish.

Why AI skips languages like Yiddish

Modern AI models learn language in a crude but effective way: they read a gigantic chunk of the internet. Trillions of words. Everything.

That works beautifully if your language has a mountain of clean, native text online. English has the whole web. Yiddish has... something much smaller, and here's the trap: a lot of what is online in Yiddish isn't really written by Yiddish speakers. It's machine-translated from other languages — the output of the very tools that mangle the language. Feed that to an AI and it learns Yiddish from counterfeit Yiddish.

Imagine trying to learn to cook from a library of microwave dinners. The labels say "food." You'd end up with very strange ideas about what dinner is.

That's the state most smaller languages are in. Not because people stopped speaking them — Yiddish has millions of speakers and centuries of world-class literature — but because that richness lives in books, classrooms, kitchens, and conversation. Not on the web.

What the team actually did

The fix is conceptually simple and heroic in execution. Three steps.

First, build a real pantry. The researchers assembled a dataset they named Oytser — Yiddish for "treasure" — of genuinely native Yiddish: today's Yiddish websites and publications, plus a large collection of literary works. Real sentences written by real people for other real people.

Second, give an existing model a Yiddish education. They took Llama — Meta's openly available model — and had it study the treasure. Not rebuild AI from scratch; teach an already-smart student one more language, properly, from native sources this time.

Third, write an exam. To prove the education worked — not just to them, to everyone — they built Kashes ("questions"), a test covering Yiddish translation, grammar, and understanding. Then they sat all the models down for it.

The report card

The results were not close. On the exam, the Yiddish-educated model beat the same-size versions of the big general models — including its own un-tutored base — on essentially every Yiddish task.

In human terms:

  • Translation into Yiddish went from "embarrassing at parties" to genuinely usable. The kinds of texts you could now machine-translate and then lightly fix, instead of rewriting from scratch.
  • Understanding of word forms and grammar roughly doubled on some tests. Yiddish is a language full of variant spellings and rich word-building; this is exactly where the general models fell down and where MameLoshnLM leapt ahead.
  • The paper's own summary: the model better captures the lexical and morphological patterns that define the language — the idioms, the style, the cultural voice. The counterfeit-Yiddish problem was, largely, un-counterfeited.

Why this is bigger than Yiddish

Here's the part that should interest you even if you've never heard a word of Yiddish.

This project is a recipe. Collect native material. Teach an existing open model. Write an exam to prove it. The whole thing was done by a small team with open tools and published for anyone to repeat.

That recipe works for Ladino. For Judeo-Arabic. For the hundreds of languages the big AI labs will never prioritize, because there's no market of a hundred million customers. AI doesn't have to flatten the world's cultures into one global English-ish mush — but it only preserves what it's fed. Someone has to do the feeding.

Yiddish just showed how.

The fine print (because there's always fine print)

  • The model learned partly from old scanned books, so it inherited some typos and old-fashioned spelling quirks — like a student who studied in a beautiful but slightly crumbling library.
  • It's released for research and education under a non-commercial license. A foundation, not a product.
  • It's a base model — a brilliant student of the language, not yet a chatty assistant. Building the conversational layer is the next step.

So what?

If you teach Yiddish, learn Yiddish, or archive Yiddish: the tools around you just got dramatically better. Translation, transcription cleanup, search across a century of literature — all of it now has a real engine underneath.

And if you just love watching how AI actually develops: skip the hype cycle for a week and read this paper instead. The most interesting thing happening in AI right now isn't bigger models. It's people using the big models as raw material and teaching them the things nobody else will.

Which language should get this treatment next?


Links: the model is here, the paper is here, and the benchmark materials are in the same Yiddish-NLP organization.