Track new features, model updates, and breaking changes across Sarvam AI’s APIs and products. Use the filters below to narrow the timeline to a single product - Model APIs, Voice Agents, Doc Agents, Content Studio, or Cowork.
Subscribe to updates: append .rss, .atom, or .json to this page’s URL - e.g. https://docs.sarvam.ai/changelog.rss.
Bulbul V4 for Voice Agents
Bulbul V4 is now available for Voice Agents, alongside Bulbul V3. The voice picker is now a filterable catalog: browse and preview voices, filter the list to find the right fit, and choose between Bulbul V3 and Bulbul V4 per agent. See Speakers & voice.
New Voice Names in Bulbul V4
Voices in Bulbul V4 carry new, more descriptive display names: Shubh is now Shubh - Hinglish Ad Voice and Ritu is now Ritu - Hindi Support Agent. The full V3-to-V4 name mapping is in Speakers & voice.
New and Improved Extract
Extract is now more accurate across Indic, English, and mixed-language documents, with lower latency and support for longer documents. Every extracted value now includes a citation to its source, viewable in the UI, so each field can be verified at a glance. Extract rate limits have also been increased. Get started with the Document AI overview.
Voice Cloning Streaming
Voice cloning now supports streaming audio. Use POST /voices/clone/stream for a single binary HTTP stream, or GET /voices/clone/ws for a persistent WebSocket that accepts incremental text. See the HTTP streaming and WebSocket guides, or the API Reference.
Dubbing API Now Available
The Dubbing API is now available for programmatically localizing audio and video while preserving each speaker’s voice. Learn how to create and manage dubbing jobs in the Dubbing API documentation, or explore the endpoints in the API Reference.
Voice Cloning API Now Available
The Voice Cloning API is now available for generating speech in a cloned voice from a short reference recording. Get started with the Voice Cloning API documentation, or explore the endpoints in the API Reference.
Invite Teammates from Workspace Settings
Workspace owners can now invite people directly from the workspace members page, with the workspace pre-selected, instead of going through organisation-level settings.
Slack Connect Made Simple
Connecting an agent to Slack no longer means creating a Slack app by hand and pasting secrets into a form. Copy one command into your Slack workspace, approve the permissions, and you are done - no Slack dashboard required.
Live Voice Input
Added live voice input: words appear on screen while you are still speaking, instead of recording a clip and waiting for it to be transcribed. The 30-second limit is gone; you can speak for up to five minutes in one go.
Share Files by Link
You can now share any file your agent creates with a link - from the chat card, the Files panel, or the preview. Choose who can open it (your workspace, your organisation, or anyone with the link) and change the audience later without breaking the link. Recipients see a clean page with just the file, and public links need no account.
Faster, Complete Outputs on Big Tasks
Agents now finish multi-step jobs faster by batching their work into a single pass instead of one step at a time. Long results from sub-tasks, such as a full report or a long rewrite, now arrive complete instead of being cut off and redone from scratch.
Keyterm Prompting on Streaming WebSockets
keyterms is now supported on the streaming WebSocket endpoints — /speech-to-text/ws (legacy) and /speech-to-text-realtime/ws (realtime) — with model="saaras:v4", matching the existing REST and Batch behaviour. Pass a JSON-encoded list of up to 50 terms (64 characters each) as the keyterms query parameter when opening the connection. See Keyterm Prompting.
Scheduled Agents Deliver Their Files
Fixed an issue where agents on a schedule, such as a daily report, posted their text but dropped the files they created. Files now arrive with the message, every time. See Scheduled Tasks.
Edit Parts of an Artifact by Commenting
You can now change a specific part of an artifact your agent built by commenting on it directly. Turn on Comment in the preview, click the block or button you want changed, and type your feedback. Each comment attaches to that element and reaches the agent with your next message, so you no longer have to describe where the problem is.
A Richer Message Composer
The message box now has two quick commands: @ to mention an agent, and / to pull in skills, knowledge bases, and connectors. What you pick is attached right where you are typing, so the agent gets what it needs without you configuring it elsewhere.
Tests for Voice Agents
Agents can now be regression-tested with Tests: describe the conversations that matter as test cases, and the platform plays a simulated user against a committed version of your agent while an AI judge grades each expected behavior Pass or Fail. Suites carry shared variables and global guardrails, runs execute each case up to 20 times to expose flakiness, and every execution returns the full transcript with per-behavior reasoning. The same suites can be created, run, and read over REST — see the Tests API overview and Best practices.
Voice Agents MCP Server
Voice agents are now available over MCP, so Claude, Cursor, VS Code, and other MCP clients can build, deploy, test, and analyse voice agents directly from your assistant. It’s a hosted, remote server — nothing to install — reachable over streamable HTTP at https://mcp.sarvam.ai/voice-agents and authenticated through Sarvam SSO (OAuth 2.1). See the MCP Server page for setup, the full tool surface, and security guidance.
GLM-5.3
Added GLM-5.3 (glm5.3) as an open-weight chat model on /v2/chat/completions. It offers a 1,048,576-token context window with tool calling and visible reasoning, priced at ₹126 / ₹23.4 / ₹396 per 1M tokens (input / cached input / output), with reasoning billed as output. Beta access is granted per API key; contact us to request whitelisting.
Boards API
Boards can now be created, edited, and run over REST instead of only in the dashboard. Discover the tables available to your workspace with query-schema, develop SQL against query-preview, then save it as a widget and run a whole tab in one call. Recurring delivery of a widget’s result to email or Slack is configurable through notification rules. Authenticate with X-API-Key as with the rest of the Voice Agents API. See the Boards API overview, Writing board SQL, and Best practices.
Data Retention Controls
Control how long Sarvam keeps your data, self-serve. Settings → Workspace → Data retention now lets an org Owner set a workspace-wide default retention period, plus a per-product override for Model APIs and Voice Agents — support for the rest of the product line is rolling out. Changing the period only affects data going forward; nothing already stored is deleted early or extended retroactively. See Data retention.