Analytics Vidhya

Bodhan AI Launches Four Indic Models for OCR, Speech and Translation

Bodhan AI Launches Four Indic Models for OCR, Speech and Translation

Analytics Vidhya

Bodhan AI, in collaboration with AI4Bharat, has released four specialized models targeting document parsing, translation, speech recognition and speech generation across Indian languages. The suite includes *IndicOCR*, *Indic-Translate*, *Indic-Transcribe* and *Indic-Speak*, each addressing a distinct challenge in making Indian-language content machine-readable and accessible.

*IndicOCR* parses printed text in all 22 scheduled languages and handwriting in 12, preserving tables and equations. *Indic-Translate*, a fine-tune of Gemma 4, focuses on document-level translation that retains formatting like Markdown and LaTeX while supporting code-mixed and Romanized input. *Indic-Transcribe* offers two variants, Core and Flex, balancing native-script accuracy against flexible Romanized or mixed-script transcription. *Indic-Speak* rounds out the suite with text-to-speech across 45 voices and 12 scripts, capable of reading code-mixed sentences without explicit language tags.

The models are available via a hosted API and Hugging Face weights, with per-model pricing. While benchmarks show strong performance gains over prior tools like IndicTrans2, the release notes several limitations, including incomplete handwriting coverage, English-pivoted translation between Indian languages, and audio-length restrictions for transcription.