Sarvam-30B Deprecated
sarvam-30b has been fully deprecated and moved to the Deprecated Models section, and is no longer available through the Chat Completions API. Sarvam-105B is now the only supported Sarvam chat model - migrate any remaining sarvam-30b integrations to sarvam-105b.
Saaras v4 for Speech to Text
Introduced saaras:v4, available alongside saaras:v3 across the REST, WebSocket, and Batch Speech-to-Text endpoints. It supports the same five output modes (transcribe, translate, verbatim, translit, codemix) and adds Global English alongside Indian English, on top of the existing 22 Indic languages. saaras:v3 remains the default, recommended model.
Document AI Launched
Replaced the Document Digitization API with Document AI (/doc-ai/v1), powered by an upgraded Sarvam Vision 1.5 trained for both OCR and key-value extraction. Two endpoints: Digitise for full-document OCR with layout and tables preserved, and Extract for schema-based field extraction to JSON, CSV, or XLSX. The earlier Document Digitization API is now legacy.
Telephony, Frameworks & Migrations
Added integration guides for Twilio, Plivo, and Exotel telephony, plus Vapi and LangChain; LiveKit and Pipecat production-tuning guides for Voice Agents; and migration guides for developers switching from Cartesia, Deepgram, ElevenLabs, or Gemini.
Realtime API for Live Transcription
Introduced a new Realtime API - saaras:v3-realtime over WebSocket - purpose-built for voice agents and live transcription. It adds true partial (interim) transcripts as the user speaks, live mid-stream reconfiguration, and millisecond-based VAD tuning. This supersedes the legacy streaming WebSocket for new voice-agent and live-transcription work; the older saaras:v3 streaming WebSocket remains generally available and is now documented under Legacy.
GLM-5.3 and Gemma 4 31B (Beta)
Introduced two open-weight chat models served on Sarvam infrastructure via the new /v2/chat/completions endpoint. GLM-5.3 (glm5.3) offers a 1,048,576-token context window with tool calling and visible reasoning, for tasks that exceed Sarvam’s own 128K context. Gemma 4 31B (gemma4) adds image input alongside tool calling. Neither model is tuned for Indian languages - for Indian-language workloads, continue using Sarvam-105B or Sarvam Vision. Beta access is granted per API key; contact us to request whitelisting.
Self-Hosted Deployments on AWS
Sarvam models can now run inside your own AWS account via Self-Hosted Deployments. Subscribe to a model package on the AWS Marketplace and deploy Saaras v3, Bulbul v3, or Sarvam Vision as an Amazon SageMaker endpoint in your own VPC - same models as the Managed API, but your audio and documents never leave your infrastructure.
LLM Pricing Updated
sarvam-105b pricing increased to ₹29.28 / ₹10.98 / ₹73.2 per 1M tokens (input / cached input / output), up from ₹4 / ₹2.5 / ₹16, effective August 5, 2026. sarvam-105b-conversations is priced the same as sarvam-105b. Pricing was also published for the beta open-weight models - GLM-5.3 at ₹126 / ₹23.4 / ₹396 and Gemma 4 31B at ₹36.6 / ₹13.73 / ₹91.5 - with GLM-5.3’s reasoning tokens billed as output. See the pricing page for details.
Conversations is Now Voice Agents
The product formerly called Conversations (internally, “Samvaad”) is now Voice Agents, an end-to-end stack for building, deploying, and monitoring voice agents on telephony. References to platform.sarvam.ai were updated accordingly.
Vobiz Telephony Integration
Added a Vobiz telephony integration guide for Voice Agents.