Advanced NLP Interview Questions, issue 9, Dec 15, 2025
The Tokenization Trap in Semitic Languages
Senior NLP Engineer interview at Google DeepMind, and the interviewer asks:
“We need to adapt our English-centric LLM to support Arabic and Hebrew. How do you adjust the tokenizer?”
Your tokenizer works great in English - until Arabic and Hebrew expose the hidden assumptions baked into subword models.
The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.