Two things Sarvam AI did in the last week of September 2026 show how the company has changed. It launched Saaras V4, a new speech-recognition model for 22 Indian languages and English. And it was reported to be reconfiguring its platform to host third-party and US-built models alongside its own — a step away from being "the Indian-language model company" towards being an India-hosted AI platform.
For a business deciding what to build on, the useful question isn't whether Sarvam is an important company (it clearly is), but which parts of its lineup are ready to use today, what they cost, and when they beat the global alternatives. That's what this post covers, from Sarvam's own documentation and dated reporting.
How it got here
Sarvam released Sarvam-1 in October 2024 and the 24-billion-parameter Sarvam-M in May 2025. In April 2025, MeitY selected it to build a sovereign foundation model under the IndiaAI Mission, with access to government-supported GPU compute. At the India AI Impact Summit in New Delhi in February 2026 it unveiled two larger models — Sarvam-30B and Sarvam-105B, both mixture-of-experts designs, the larger with a 128K-token context window — and launched Indus, a consumer chat app, on 20 February 2026. In June 2026 it raised a $234 million Series B led by HCLTech at a $1.5 billion valuation.
Saaras V4: the speech model
Saaras V4 is an automatic speech-recognition model that covers 22 Indian languages plus English, detects the spoken language automatically and handles code-mixed speech — the Hinglish-style switching that trips up many global models. It can output five formats: verbatim transcripts, normalised text, code-mixed text, transliteration and translation. Sarvam says streaming latency is below 150 milliseconds and reports a language-identification error rate of 5.22% across the 22 languages, or 2.9% across the top ten.
Two cautions. Those accuracy numbers are Sarvam's own; no independent benchmark had been published with the launch. And the launch coverage didn't include V4-specific pricing. Sarvam's pricing page currently lists speech-to-text at ₹30 per hour of audio, or ₹45 per hour with speaker diarisation. The model is available by API with Python and Node.js SDKs, and integrates with voice-agent frameworks including LiveKit Agents and Pipecat.
The platform shift: other people's models, on Sarvam
The more strategic change is on the platform side. Sarvam's own API pricing page already lists third-party models next to its own: alongside Sarvam 105B (₹29.28 per million input tokens and ₹73.20 per million output tokens), it offers Google's Gemma 4 31B, Zhipu's GLM-5.3 and DeepSeek V4 Flash, all billed in rupees. Reporting on 26 September 2026 described Sarvam reworking its infrastructure to host US-built models too, aimed at enterprises that want one Indian vendor and data kept locally or on private networks.
For Indian businesses, that's a meaningful option. It could let a company use a widely known open model while paying in rupees, getting an Indian invoice, and keeping inference with an Indian vendor rather than calling an overseas API. Sarvam markets its inference platform as India-hosted, but where a particular model runs is a contract question — confirm the region for the specific model you use rather than assuming it applies to the whole catalogue.
The rest of the lineup, and what it costs
Text-to-speech (Bulbul) is priced at ₹30 per 10,000 characters, as is voice cloning. Translation and transliteration cost ₹20 per 10,000 characters; document translation is ₹5 per 1,000 billable characters; document digitisation is ₹0.50 a page. New accounts get ₹100 of free credit, which is enough to test the speech and translation APIs on real samples.
Samvaad, Sarvam's voice-agent platform for customer conversations, has handled 325 million minutes of calls and was reported by Inc42 in August 2026 at ₹3.5 per minute — well under the ₹8–12 range that article quoted for the wider industry. The same report listed newer products including Sarvam Vision 2.0 for documents, Sarvam Work for workplace agents, a desktop dictation app, and Sarvam Kaze smart glasses.
When to pick Sarvam, and when not to
It's a strong first choice for Indian-language speech — transcription, voice bots, dubbing — where accents, code-mixing and language coverage matter more than raw reasoning. It also makes sense where you want an Indian vendor, rupee billing and a GST invoice, or local data handling, for procurement or regulatory reasons — Sarvam passes all seven India-readiness checks in our catalogue, though, like every listing, those details are editor-compiled and worth confirming with the vendor.
It's a weaker choice for English-first work that depends on top-tier reasoning, long agent workflows or a large ecosystem of plug-ins and integrations, where ChatGPT, Claude and Gemini still lead. Many teams will end up using both: a global model for general reasoning, and Sarvam for the Indian-language voice and text layer that the global models handle less well.
What to watch
The two open questions are independent benchmarks for Saaras V4 against global speech models on Indian audio, and firm details on which third-party models Sarvam will host, in which regions, and on what terms. We'll update this post and Sarvam's tool page as those land. Prices above are from Sarvam's pricing page as checked on 28 September 2026 and change often — confirm before you budget.