Share

Bengali Speech Recognition Fixed by Tokenizer Swap

Bengali Speech Recognition Fixed by Tokenizer Swap

The Cryptonomist

The Cryptonomist

Researchers Sanjid Hasan and Md. Abdur Rahman have identified why compact speech recognition models fail catastrophically with Bengali and developed a surprisingly efficient fix. The problem lies not in the model architecture itself but in English-centric byte-level tokenizers that fragment Bengali words into unstable token chains, triggering what they call *autoregressive collapse* during inference.

Their solution, *vocabulary transplantation*, replaces the decoder vocabulary with BanglaBERT WordPiece tokens designed for Bengali's morphological richness and native script. This surgical intervention requires no costly pre-training from scratch. Results were dramatic: token fertility dropped from 9.16 to 1.30, and autoregressive sequence length fell by 85.8%.

Testing on the 882-hour Lipi-Ghor dataset yielded a 21.54% Word Error Rate and a Real-Time Factor of 0.0053, proving the approach enables both accurate and extremely fast edge deployment. Accepted at ICML 2026's MusIML Workshop, the research offers a scalable blueprint for adapting lightweight ASR systems to other morphologically complex, non-Latin languages historically underserved by English-centric AI infrastructure.