Advanced NLP Interview Questions, issue 9, Dec 15, 2025

The Tokenization Trap in Semitic Languages

Senior NLP Engineer interview at Google DeepMind, and the interviewer asks:

We need to adapt our English-centric LLM to support Arabic and Hebrew. How do you adjust the tokenizer?

Your tokenizer works great in English - until Arabic and Hebrew expose the hidden assumptions baked into subword models.

The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.

Read it on Substack

Get the next one

Free on Substack. Unsubscribe in one click.