Share

IISc Builds Atypical Speech Dataset for Indian Languages

IISc Builds Atypical Speech Dataset for Indian Languages

The Hindu

The Hindu

Researchers at the Indian Institute of Science, Bengaluru, are addressing a critical gap in speech technology by building the *Vaani Atypical Speech Corpus*, India's first dataset of atypical speech patterns in Indian languages. The pilot initiative has collected approximately 10 hours of atypical speech from around 40 speakers across three languages.

Atypical speech, which deviates from conventional patterns due to neurological, cognitive, or motor conditions, has been largely overlooked in Indian language speech recognition systems. While global efforts exist to train ASR systems on atypical speech, no such data existed for Indian languages until now.

The initiative is part of the larger *Project Vaani*, which aims to capture 150,000 hours of natural speech data from 1 million people across nearly 800 Indian districts. So far, the project has covered 165 districts, recorded over 31,000 hours of speech data across 105 languages, and transcribed more than 2,000 hours. The team uses image-prompted recordings to collect multi-modal data in local dialects, ensuring diversity across gender, age, and socio-economic backgrounds.