SraVaani-1.0 covers 65 Indian languages and dialects, with the released checkpoint listing 63.
The model shows strong results across low-resource, tribal and dialect languages.
SraVaani is open source and available through Hugging Face for further development.
India has taken a major step in speech technology with SraVaani-1.0, an open-source automatic speech recognition model from SPIRE Lab at IISc and ARTPARK@IISc. The new system can handle 65 Indian languages and dialects, with a strong focus on languages that receive little support from major speech AI systems. The SraVaani-1.0 research paper appeared on arXiv on August 8, 2026.
The project stands out for its focus on India's long list of low-resource languages. Large commercial speech systems often perform best on languages with huge amounts of digital speech data. Many Indian languages and local dialects do not have that advantage. SraVaani takes a different route. Its creators built the system around a large Indian speech corpus and added methods that help the model learn useful links between speech, language and visual information.
SraVaani starts with the VAANI corpus, which contains a much wider range of Indian languages. The first stage used 31,255 hours of unlabelled speech for self-supervised speech pretraining. VAANI itself covers 105 languages across its broader corpus. This large base gives the model exposure to speech patterns that standard datasets often miss.
The next stage adds an unusual feature. SraVaani uses about 11.85 million paired audio-image samples to align speech with visual information. This approach helps the speech encoder connect spoken words with meaning and context. The idea matters for languages with limited labelled speech, where extra forms of useful information can help improve recognition.
The final stage uses about 31,263 hours of labelled Indian speech from 24 public datasets. This stage covers the reported 65 languages and dialects. Together, these stages give SraVaani a broad base for speech recognition across India's diverse language landscape.
The research paper reports coverage across 65 Indian languages and dialects. The current Hugging Face model card, however, describes the released SraVaani-1.0 checkpoint as supporting 63 languages and dialects. This small difference matters when the project gets described in technical reports. The safest description remains that the SraVaani-1.0 research system covers 65 languages and dialects, while the current released checkpoint lists 63.
SraVaani uses the FastConformer architecture with a hybrid TDT-CTC decoder. The released model has about 430 million parameters and takes roughly 900 MB in FP16. These figures place the system within a practical range for developers who need a capable multilingual speech model rather than a very large general AI model.
Also Read - Top Multilingual Text-to-Speech Tools in 2026
The most important result does not come from the language count alone. SraVaani shows particular strength in languages that lack strong speech AI support.
A recent ARTPARK evaluation compared SraVaani with Google Gemini 3 Flash, Sarvam Saaras v3 and IndicConformer-600M-Multilingual across eight benchmarks. SraVaani recorded the lowest word error rate on many language-dataset combinations. The results also showed strong performance on several low-resource languages and dialects.
Of 17 languages where a direct comparison with IndicConformer proved possible, SraVaani recorded a mean word error rate of 28.4%, compared with 30.2% for IndicConformer. A lower word error rate means fewer mistakes in the final transcript.
The results become more notable for tribal and dialect languages. ARTPARK reports that SraVaani produced output for 51 tribal and dialect languages in the relevant VAANI benchmark region where competing systems either produced no output or recorded word error rates above 80%.
SraVaani has also moved beyond a research paper. SraVaani-1.0 is available on Hugging Face, and its model repository lists an MIT license. Access to the model files currently requires acceptance of Hugging Face conditions. ARTPARK also provides a live SraVaani demo and continues to update its model collection.
Recent project activity shows a fast development cycle. The technical paper arrived on August 8, followed by an ARTPARK article on August 10 that explained the role of vision and audio alignment. On August 11, ARTPARK published a guide for fine-tuning SraVaani with custom speech data. The organisation's Hugging Face page also shows recent updates to SraVaani-1.0 and a SraVaani-0.5-live model for real-time speech-to-text use.
Also Read - Top AI Voice Generator and Text-to-Speech Platforms
SraVaani's real value lies in its reach. India's speech technology cannot serve the whole country if strong performance remains limited to a small group of major languages. A system that can recognise tribal languages, regional dialects and other low-resource forms of speech can expand access to voice-based technology.
The model could support local-language government services, education tools, accessibility products, healthcare interfaces and voice assistants. Its open availability also gives researchers and developers a foundation for new applications and further language work.
SraVaani therefore represents more than a 65-language milestone. It shows a clear attempt to build speech AI around India's full linguistic diversity, rather than only its largest digital languages. The latest results suggest that this approach can produce useful gains precisely where mainstream speech systems often struggle most.
1. What is SraVaani?
SraVaani is an open-source automatic speech recognition model developed by SPIRE Lab at IISc and ARTPARK@IISc.
2. How many Indian languages does SraVaani support?
The research paper reports coverage across 65 Indian languages and dialects, while the current released checkpoint lists 63.
3. What makes SraVaani different from other speech models?
SraVaani places strong emphasis on low-resource, tribal and dialect languages that often have limited support from mainstream speech systems.
4. Is SraVaani available to developers?
Yes. SraVaani-1.0 is available on Hugging Face under an MIT license, with access to model files subject to Hugging Face conditions.
5. How accurate is SraVaani?
Across 17 languages with a direct comparison, SraVaani recorded a mean word error rate of 28.4%, compared with 30.2% for IndicConformer.