Why Voice Agents
If you are building a voice agent for customers in India, Sarvam Voice Agents gives you the whole stack in one platform: speech recognition, reasoning, and speech synthesis built for Indian languages, plus telephony, testing, outbound campaigns, and analytics. Pricing is per minute, from ₹3.50 on pay as you go down to ₹3 on higher plans.
What you get in one platform
- 11 Indian languages plus English. Hindi, Tamil, Telugu, Bengali, Marathi, Gujarati, Kannada, Malayalam, Punjabi, Odia, and Assamese, including code-mixed speech, alphanumerics, and proper nouns. See Overview.
- Sarvam’s own models, self-hosted by Sarvam. Saaras v4 for speech recognition, Bulbul v3 and v4 for speech synthesis, and
sarvam-105b-conversationsfor real-time dialogue, orchestrated for you. See Models. - Data residency in India. All data, including any PII, stays within India.
- A run-time built for real phone calls. Interruptions and barge-in, language switching mid-call, voicemail detection, hold, long silences, and correct pronunciation of names and numbers are handled out of the box. See Run-time.
- Telephony included. Rent an Indian number from Sarvam or bring your own provider. The same agent also runs on WhatsApp, web, and API.
- Test before you go live. Run simulated test calls against scripted scenarios.
- Outbound at scale and analytics. Run campaigns, then read call logs, analytics, and custom boards.
- Build from your AI assistant. Connect Claude, Cursor, or VS Code over the MCP server to build, test, and deploy agents.
Build it yourself vs Sarvam Voice Agents
A production voice agent needs more than a speech API. Here is what you would assemble yourself, and what Voice Agents already includes.
What it costs
Self-serve plans are prepaid from your credit balance. Enterprise is postpaid and invoiced monthly. For example, 10,000 minutes a month costs ₹35,000 on pay as you go. Telephony is billed separately.
When to use the APIs instead
Use Sarvam’s speech-to-text, text-to-speech, and chat APIs directly if you already run your own orchestration (for example LiveKit or Pipecat) and want full control of every component. For everything else, start with Voice Agents.