Sarvam's Saaras V4 transcribes 22 Indian languages at ₹30 an hour

single source· 1 articles · confidence: medium · first seen 2026-09-26 21:56 UTC

What this means for you

Worth a test if you transcribe Indian-language audio: ₹30 an hour, streaming output, key-term prompting and five output formats from one model. No accuracy figures or independent evaluation have been published, so measure it on your own audio before you switch anything over.

Sarvam AI has released Saaras V4, a speech-to-text model that covers all 22 Indian languages plus global English. It pairs an audio encoder with a 3-billion-parameter decoder built from a mix of attention and state-space layers, a layer type that handles long inputs more cheaply than attention alone. It takes up to 50 key terms as prompts to steer transcription, offers five output formats from one model, and streams with first-token latency under 150 milliseconds. It is available now through Sarvam's API at ₹30 per hour. No accuracy figures, evaluation date or harness accompany the release; the performance claims are the company's.

Key facts

  • ·Saaras V4 covers all 22 Indian languages plus global English, according to Sarvam AI. source
  • ·The model pairs an audio encoder with a 3B hybrid state-space decoder. source
  • ·It supports keyterm prompting for up to 50 terms. source
  • ·It offers five output modes from a single model, with streaming first-token latency under 150 ms. source
  • ·It is available through Sarvam's API at ₹30 per hour. source

What the sources say

  • MarkTechPost — Relays the vendor's specification sheet: language coverage, decoder size, streaming latency and hourly API price.

Sources

The original reporting. Follow these — they did the work.

← the wire