> For clean Markdown of any page, append `.md` to the page URL. > For a complete documentation index, see https://docs.sarvam.ai/llms.txt. > For full documentation content in one file, see https://docs.sarvam.ai/llms-full.txt. > For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.sarvam.ai/_mcp/server. # Model APIs ## Docs - [Welcome to Sarvam Model APIs Documentation](https://docs.sarvam.ai/api/getting-started/welcome.md): Welcome to Sarvam AI documentation. Access comprehensive guides, API references, quickstart tutorials, and community resources for Indian language AI development. - [Developer Quickstart](https://docs.sarvam.ai/api/getting-started/quickstart.md): Learn how to make your first API request with Sarvam AI in under 5 minutes. Complete guide with code examples for chat completion, speech-to-text, and translation APIs. - [Libraries & SDKs](https://docs.sarvam.ai/api/getting-started/sdks.md): Official Python and JavaScript clients for the Sarvam AI API: with async, retries, timeouts, streaming, and typed errors. - [Building for Indian Languages](https://docs.sarvam.ai/api/getting-started/building-for-india.md): A practical guide to shipping Indian-language AI, speech-to-text modes (code-mix, transliteration), natural Indian voices, 8kHz telephony audio, pronunciation control, and document digitization across 23 languages (22 Indian + English). - [Models](https://docs.sarvam.ai/api/getting-started/models.md): Complete overview of Sarvam AI's specialized models for Indian languages. Choose the right model for your use case - from speech processing to text generation, translation, and document intelligence. - [Saaras](https://docs.sarvam.ai/api/getting-started/models/saaras.md): Saaras v3 and v4 - Domain-aware speech translation models that convert speech directly to English text with enhanced telephony support and intelligent entity preservation. - [Bulbul](https://docs.sarvam.ai/api/getting-started/models/bulbul.md): Bulbul v3 - High-quality multilingual text-to-speech model for Indian languages with natural prosody and 30+ speaker voices. - [Sarvam-105B](https://docs.sarvam.ai/api/getting-started/models/sarvam-105b.md): Sarvam-105B - 105B parameter flagship multilingual language model delivering state-of-the-art performance on Indian language understanding, reasoning, and generation tasks. - [Open-Weight Models](https://docs.sarvam.ai/api/getting-started/models/openweight.md): Open-weight models served on Sarvam infrastructure through the OpenAI-compatible chat completions API, using your existing Sarvam API key and credits. - [GLM-5.3](https://docs.sarvam.ai/api/getting-started/models/openweight/glm-5-3.md): GLM-5.3 is a large-scale reasoning model built for complex software engineering and long-horizon agent tasks. It supports text input and output with a 1,048,576-token context window. - [Gemma 4 31B](https://docs.sarvam.ai/api/getting-started/models/openweight/gemma-4-31b.md): Gemma 4 31B is a 31B-parameter instruction-tuned model for text and image understanding. It supports a 131,072-token context window and tool calling. - [DeepSeek V4 Flash](https://docs.sarvam.ai/api/getting-started/models/openweight/deepseek-v4-flash.md): DeepSeek V4 Flash is a general-purpose open-weight reasoning model with text input and output, a 1,048,576-token context window, and tool calling. - [Mayura](https://docs.sarvam.ai/api/getting-started/models/mayura.md): Mayura - Advanced multilingual translation model for Indian languages with customizable translation styles, script control, and intelligent code-mixed content handling. - [Sarvam Translate](https://docs.sarvam.ai/api/getting-started/models/sarvam-translate.md): Sarvam Translate - Comprehensive translation model supporting all 22 official Indian languages with formal translation style and structured text optimization. - [Sarvam Vision](https://docs.sarvam.ai/api/getting-started/models/sarvam-vision.md): Sarvam Vision - A 3B parameter multimodal model delivering world-class Document Intelligence and visual understanding with unmatched accuracy for 23 languages (22 Indian + English). - [Sarvam-30B (Deprecated)](https://docs.sarvam.ai/api/getting-started/models/sarvam-30b.md): Sarvam-30B - 30B parameter multilingual language model optimized for Indian languages with strong reasoning, coding, and conversational capabilities. Deprecated; migrate to Sarvam-105B. - [Sarvam-M (Deprecated)](https://docs.sarvam.ai/api/getting-started/models/sarvam-m.md): Sarvam-M - deprecated 24B parameter multilingual, hybrid-reasoning language model with strong Indian language benchmark performance. - [Saarika](https://docs.sarvam.ai/api/getting-started/models/saarika.md): Saarika v2.5 - High-accuracy speech recognition model for Indian languages with superior multi-speaker handling, telephony optimization, and automatic code-mixing support. - [Credits & Rate Limits](https://docs.sarvam.ai/api/getting-started/ratelimits.md): Understand Sarvam AI rate limits by plan tier, per-API concurrency limits, and how to handle 429 and 503 errors gracefully. View your current limits on the dashboard. - [Commercial Licensing](https://docs.sarvam.ai/api/getting-started/commercial-licensing.md): Commercial production rights for Sarvam AI Output, including Bulbul v3 text-to-speech, voice cloning, and Content Studio exports. Covers credits, the Production License, and retained rights after you stop using the service. - [Errors & Troubleshooting](https://docs.sarvam.ai/api/getting-started/errors-troubleshooting.md): Central reference for Sarvam API error codes, HTTP status handling, SDK exceptions, retries, and common integration pitfalls. - [Talk to us](https://docs.sarvam.ai/api/getting-started/help.md): Get help and support for Sarvam AI APIs. Contact our team via email for technical questions, bug reports, feature requests, and enterprise inquiries, or join our Discord to connect with the builder community. - [Pricing](https://docs.sarvam.ai/api/getting-started/pricing.md): Transparent pricing for all Sarvam AI services. View rates for language models, speech-to-text, text-to-speech, text processing, and document intelligence APIs in Indian Rupees. - [Organisations & workspaces](https://docs.sarvam.ai/api/platform/organisations-and-workspaces.md): How Sarvam accounts are structured. An organisation is your billing and identity boundary; workspaces sit inside it and hold your API keys. Learn what lives where, the Owner and Member roles, and account limits. - [Create an organisation](https://docs.sarvam.ai/api/platform/create-an-organisation.md): Step-by-step guide to creating a new Sarvam organisation from the dashboard, when a separate organisation is the right choice, and what does not carry over from your existing account. - [Create a workspace](https://docs.sarvam.ai/api/platform/create-a-workspace.md): Step-by-step guide to creating a workspace inside your Sarvam organisation, how workspaces separate API keys and usage while sharing one credit balance, and the 100-workspace limit. - [Invite your team](https://docs.sarvam.ai/api/platform/invite-your-team.md): How to invite users to your Sarvam organisation, grant workspace access, change someone's role, revoke a pending invitation, and remove a member. Only organisation Owners can invite. - [Billing & payments](https://docs.sarvam.ai/api/platform/billing.md): Billing on Sarvam is organisation-level. One credit balance is shared by every workspace, API key, and product. Learn how to add credits, set up auto-recharge, and read your payment history. - [Platform FAQ](https://docs.sarvam.ai/api/platform/faq.md): Short answers on organisations vs workspaces, who can invite users and pay, where credits live, API key limits, switching organisations, and what happens when someone leaves your team. - [Chat Completion API](https://docs.sarvam.ai/api/api-guides-tutorials/chat-completion/overview.md): Get started with Sarvam AI LLM models for conversational AI. Build intelligent chat applications with native Indian language support and deep contextual reasoning capabilities. - [Responses API](https://docs.sarvam.ai/api/api-guides-tutorials/chat-completion/responses-api.md): Create stateless text, reasoning, structured-output, and tool-calling responses with Sarvam's OpenAI-compatible Responses API. - [How to list your chat messages](https://docs.sarvam.ai/api/api-guides-tutorials/chat-completion/how-to/list-your-chat-messages.md): Defines your entire conversation. - [How to control response randomness with `temperature`](https://docs.sarvam.ai/api/api-guides-tutorials/chat-completion/how-to/control-response-randomness.md): Control how focused or varied model responses are. - [How to control response diversity with `top_p`](https://docs.sarvam.ai/api/api-guides-tutorials/chat-completion/how-to/control-response-diversity.md): Method used to generate text by limiting the possibilities of the next word - [How to adjust the model's thinking level with `reasoning_effort`](https://docs.sarvam.ai/api/api-guides-tutorials/chat-completion/how-to/adjust-the-models-thinking-level.md): controls **how much effort the model puts into reasoning. - [How to encourage new topics with `presence_penalty`](https://docs.sarvam.ai/api/api-guides-tutorials/chat-completion/how-to/encourage-new-topics-in-response.md): Helps you steer the model toward introducing new concepts or topics. - [How to reduce repetition with `frequency_penalty`](https://docs.sarvam.ai/api/api-guides-tutorials/chat-completion/how-to/reduce-repetition-words-or-phrases-in-response.md): Helps you control how often the model repeats words or phrases - [How to get repeatable results using `seed`](https://docs.sarvam.ai/api/api-guides-tutorials/chat-completion/how-to/get-repeatable-results.md): Reduce output variation with best-effort seeded sampling - [How to control the response length with `max_tokens`](https://docs.sarvam.ai/api/api-guides-tutorials/chat-completion/how-to/control-the-response-length.md): control how long the model's response can be - [How to control where the model stops using `stop`](https://docs.sarvam.ai/api/api-guides-tutorials/chat-completion/how-to/control-where-the-model-stops.md): Tell the model to **stop generating further tokens. - [Speech-to-Text APIs](https://docs.sarvam.ai/api/api-guides-tutorials/speech-to-text/overview.md): Complete overview of Sarvam AI Speech-to-Text APIs including real-time, batch, and streaming options. Process audio with the Saaras model for high-accuracy transcription. - [Which Speech-to-Text API to Use](https://docs.sarvam.ai/api/api-guides-tutorials/speech-to-text/which-api-to-use.md): Compare Sarvam's Speech-to-Text APIs, REST, WebSocket, and Batch (plus their Speech-to-Text-Translate variants), and pick the right one for your audio length, latency, and feature needs. - [Speech-to-Text Rest API](https://docs.sarvam.ai/api/api-guides-tutorials/speech-to-text/rest-api.md): Process short audio files synchronously with immediate response. Instant transcription and translation for quick audio processing with multiple format support. - [Batch Speech-to-Text API](https://docs.sarvam.ai/api/api-guides-tutorials/speech-to-text/batch-api.md): Process large audio files using synchronous or asynchronous methods. Handle up to 2-hour recordings with speaker diarization, timestamps, and advanced transcription features. - [Realtime Streaming Speech-to-Text API](https://docs.sarvam.ai/api/api-guides-tutorials/speech-to-text/realtime-streaming.md): True partial transcripts, live mid-stream reconfiguration, and millisecond-based VAD tuning with saaras:v3-realtime and saaras:v4, Sarvam's next-generation WebSocket streaming API for voice agents and live transcription. - [Best Practices for Speech-to-Text](https://docs.sarvam.ai/api/api-guides-tutorials/speech-to-text/best-practices.md): Connection handling, VAD tuning, and production recommendations for Sarvam AI's Speech-to-Text API. - [FAQs](https://docs.sarvam.ai/api/speech-to-text/faq.md): Frequently asked questions about Sarvam AI speech-to-text services. Get answers about models, pricing, language support, audio formats, and implementation best practices. - [How to select output mode](https://docs.sarvam.ai/api/api-guides-tutorials/speech-to-text/how-to/select-output-mode.md): Choose the right output mode for your speech-to-text use case with Saaras v3 and v4. - [How to specify language codes](https://docs.sarvam.ai/api/api-guides-tutorials/speech-to-text/how-to/specify-language-codes.md): Use BCP-47 language codes for accurate speech-to-text transcription with Saaras v3 and v4. - [How to enable speaker diarization](https://docs.sarvam.ai/api/api-guides-tutorials/speech-to-text/how-to/enable-speaker-diarization.md): Identify and distinguish between multiple speakers in audio using the Batch API. - [Keyterm Prompting](https://docs.sarvam.ai/api/api-guides-tutorials/speech-to-text/how-to/keyterms.md): Use keyterms to help Saaras v4 recognize important names, places, brands, and technical terms. - [IVR & Contact Center Voice Agents](https://docs.sarvam.ai/api/api-guides-tutorials/speech-to-text/use-cases/ivr-contact-center.md): Architecture, model configuration, and latency targets for building telephony-based IVR and contact-center voice agents with Sarvam AI. - [BFSI Voice Bots](https://docs.sarvam.ai/api/api-guides-tutorials/speech-to-text/use-cases/bfsi-voice-bots.md): Architecture, model configuration, and compliance guardrails for building banking, financial services, and insurance voice agents with Sarvam AI. - [EdTech Voice Tutors](https://docs.sarvam.ai/api/api-guides-tutorials/speech-to-text/use-cases/edtech.md): Architecture, model configuration, and pitfalls for building voice-based tutoring agents with Sarvam AI. - [Government Services Voice Agents](https://docs.sarvam.ai/api/api-guides-tutorials/speech-to-text/use-cases/government-services.md): Architecture, model configuration, and pitfalls for building citizen-facing government scheme and services voice agents with Sarvam AI. - [Agri & Rural Voice Agents](https://docs.sarvam.ai/api/api-guides-tutorials/speech-to-text/use-cases/agri-rural.md): Architecture, model configuration, and pitfalls for building farmer-advisory and rural voice agents with Sarvam AI, tuned for low-bandwidth, feature-phone-first users. - [Streaming Speech-to-Text API](https://docs.sarvam.ai/api/api-guides-tutorials/speech-to-text/streaming-api.md): Real-time audio transcription and translation with WebSocket connections. Low-latency streaming for live applications with instant results and interactive features. - [Text-to-Speech Overview](https://docs.sarvam.ai/api/api-guides-tutorials/text-to-speech/overview.md): Complete overview of Sarvam AI Text-to-Speech APIs using Bulbul v3 model. Convert text to natural speech with real-time and streaming options for Indian languages. - [Voices](https://docs.sarvam.ai/api/api-guides-tutorials/text-to-speech/voices.md): Listen to every Sarvam AI text-to-speech voice before you choose. Audio previews for all Bulbul v3 speakers, including multilingual samples. - [Which Text-to-Speech API to Use](https://docs.sarvam.ai/api/api-guides-tutorials/text-to-speech/which-api-to-use.md): Compare Sarvam's Text-to-Speech APIs, REST, HTTP streaming, and WebSocket, and pick the right one for your latency, text length, and interactivity needs. - [Text-to-Speech Rest API](https://docs.sarvam.ai/api/api-guides-tutorials/text-to-speech/rest-api.md): Real-time conversion of text into speech using customizable voices. Instant audio generation with multiple voice options and various audio formats for Indian languages. - [HTTP Streaming API](https://docs.sarvam.ai/api/api-guides-tutorials/text-to-speech/streaming-api/http-stream.md): Stream TTS audio over a single HTTP POST request. No WebSocket setup, no connection management, just POST text and pipe the audio response. - [Streaming Text-to-Speech API](https://docs.sarvam.ai/api/api-guides-tutorials/text-to-speech/streaming-api/web-socket.md): Real-time conversion of text into speech using WebSocket connections. Efficient streaming for long texts with progressive audio generation and low-latency playback. - [Pronunciation Dictionary](https://docs.sarvam.ai/api/api-guides-tutorials/text-to-speech/pronunciation-dictionary.md): Teach Bulbul v3 how to say specific words, brand names, abbreviations, regional terms, exactly the way you want, across all 11 supported languages. - [Best Practices for Writing Text for TTS](https://docs.sarvam.ai/api/api-guides-tutorials/text-to-speech/best-practices.md): A guide to writing text that produces natural-sounding speech output with Sarvam AI Bulbul. - [How to Set the Language](https://docs.sarvam.ai/api/api-guides-tutorials/text-to-speech/how-to/set-the-language.md): Defines the language for text normalization before speech synthesis. - [How to change the speaker voice](https://docs.sarvam.ai/api/api-guides-tutorials/text-to-speech/how-to/change-the-speaker-voice.md): Learn how to choose specific voices for text-to-speech output using the speaker parameter. Explore Bulbul v3's 30+ natural-sounding voices for different languages and use cases. - [How to adjust the pitch (tone)](https://docs.sarvam.ai/api/api-guides-tutorials/text-to-speech/how-to/adjust-the-tone.md): Control the tone of the synthesized speech (bulbul:v2 only). - [How to adjust the pace (speed)](https://docs.sarvam.ai/api/api-guides-tutorials/text-to-speech/how-to/adjust-the-speed.md): Controls the speed at which the speech is delivered. - [How to adjust the loudness (volume)](https://docs.sarvam.ai/api/api-guides-tutorials/text-to-speech/how-to/adjust-the-loudness.md): Controls the volume level of the generated audio (bulbul:v2 only). - [How to set the sample rate (audio quality)](https://docs.sarvam.ai/api/api-guides-tutorials/text-to-speech/how-to/set-the-sample-rate.md): Controls the audio quality and size of the generated output. - [How to enable text preprocessing](https://docs.sarvam.ai/api/api-guides-tutorials/text-to-speech/how-to/enable-text-preprocessing.md): improves pronunciation (bulbul:v2 only). - [How to set the audio format for output using `output_audio_codec`](https://docs.sarvam.ai/api/api-guides-tutorials/text-to-speech/how-to/set-audio-format-for-output.md): Choose the audio format for TTS streaming output. - [How to set `output_audio_bitrate`](https://docs.sarvam.ai/api/api-guides-tutorials/text-to-speech/how-to/set-bitrate-for-output.md): Control the quality and size of the synthesized audio output. - [How to set maximum length for sentence splitting using `max_chunk_length`](https://docs.sarvam.ai/api/api-guides-tutorials/text-to-speech/how-to/set-maximum-length-for-sentence-splitting.md): Control how long each sentence chunk can be when splitting text for streaming TTS. - [How to set buffer size to start processing in Streaming TTS with `min_buffer_size`](https://docs.sarvam.ai/api/api-guides-tutorials/text-to-speech/how-to/set-buffer-size-to-start-processing.md): Define when the TTS engine should start processing text from the buffer. - [Text Processing Overview](https://docs.sarvam.ai/api/api-guides-tutorials/text-processing/overview.md): Complete overview of Sarvam AI Text Processing APIs including translation, transliteration, and language identification for 22+ Indian languages using Mayura and Sarvam-Translate models. - [Text Translation API](https://docs.sarvam.ai/api/api-guides-tutorials/text-processing/translation.md): Complete overview of Sarvam AI Text Translation API supporting English to Indian languages and vice versa with multiple translation modes and high accuracy. - [Transliteration API](https://docs.sarvam.ai/api/api-guides-tutorials/text-processing/transliteration.md): Complete overview of Sarvam AI Transliteration API for script conversion between Indian languages. Convert between Roman, Devanagari, and other scripts with high accuracy. - [Language Identification API](https://docs.sarvam.ai/api/api-guides-tutorials/text-processing/language-detection.md): Identifies the language and script of input text, supporting multiple Indian languages. Automatic detection with confidence scores for multilingual text processing. - [Document Intelligence Overview](https://docs.sarvam.ai/api/api-guides-tutorials/document-intelligence/overview.md): Transform documents into structured, queryable data with Sarvam's Document AI API. Powered by Sarvam Vision 1.5 for text digitization and schema-based field extraction across 23 languages (22 Indian + English). - [Get your config_id](https://docs.sarvam.ai/api/api-guides-tutorials/document-intelligence/how-to/get-your-config-id.md): How to create a Config from an Extract project and get its config_id for use with the Document AI Extract API. - [Document Translation API Overview](https://docs.sarvam.ai/api/api-guides-tutorials/doc-translation/overview.md): Translate whole PDF, Word, spreadsheet, presentation, and web page files into 23 languages (English + 22 Indic) while preserving the original layout, with Sarvam's Document Translation API. - [How to set style guidelines](https://docs.sarvam.ai/api/api-guides-tutorials/doc-translation/how-to/style-guidelines.md): Use style_guidelines and language_specific_guidelines to control tone, terminology, and formatting in Sarvam Document Translation output. - [How to pick a genre](https://docs.sarvam.ai/api/api-guides-tutorials/doc-translation/how-to/genre-and-model-tier.md): Use genre to tune Sarvam Document Translation for your content category. - [How to use native numerals](https://docs.sarvam.ai/api/api-guides-tutorials/doc-translation/how-to/native-numerals.md): Control whether translated numbers appear as international digits (0-9) or language-specific native numerals with use_native_numerals. - [How to poll status and export results](https://docs.sarvam.ai/api/api-guides-tutorials/doc-translation/how-to/poll-and-export.md): Poll live-status for per-language translation progress, then trigger and poll an async export for each completed language's translated document from the Sarvam Document Translation API. - [FAQs](https://docs.sarvam.ai/api/api-guides-tutorials/doc-translation/faq.md): Frequently asked questions about the Sarvam AI Document Translation API. Get answers about file formats, languages, style guidelines, and job status. - [Use Cases](https://docs.sarvam.ai/api/api-guides-tutorials/doc-translation/use-cases.md): How to pick genre for common Document Translation content types. - [Translate a PDF to Hindi (or any Indic language)](https://docs.sarvam.ai/api/api-guides-tutorials/doc-translation/guides/translate-a-pdf.md): Step-by-step guide to translate a PDF with the Document Translation API — create a job, upload the file, poll progress, and download a layout-preserving translated PDF. - [Dubbing API](https://docs.sarvam.ai/api/api-guides-tutorials/dubbing/overview.md): Localize video and audio into 12 Indian languages with the Sarvam AI Dubbing API. Clone each speaker's voice across languages, control translation tone, and export video, audio, and subtitles from a single asynchronous job. - [Job Lifecycle](https://docs.sarvam.ai/api/api-guides-tutorials/dubbing/job-lifecycle.md): Every state a Sarvam AI dubbing job passes through, how to read live-status and export-status correctly, and a complete polling script that handles partial failures. - [How to specify language codes](https://docs.sarvam.ai/api/api-guides-tutorials/dubbing/how-to/specify-language-codes.md): The 12 BCP-47 language codes supported by the Sarvam Dubbing API, and how to set source and target languages for single- and multi-language jobs. - [How to control translation tone](https://docs.sarvam.ai/api/api-guides-tutorials/dubbing/how-to/control-translation-tone.md): Use the register parameter to control formality and style in Sarvam Dubbing translations: formal, common-indic, classic-colloquial, modern-colloquial, academic, or auto. - [How to choose a voice](https://docs.sarvam.ai/api/api-guides-tutorials/dubbing/how-to/choose-a-voice.md): How Sarvam AI Dubbing decides what your dub sounds like. Voice cloning preserves each original speaker across languages, with 15 preset voices as an alternative when cloning is not the right fit. - [How to choose export formats](https://docs.sarvam.ai/api/api-guides-tutorials/dubbing/how-to/choose-export-formats.md): Use export_options to auto-produce dubbed video, isolated audio, and SRT subtitles per target language, read them back from export-status, and re-export a format you did not originally request. - [FAQs](https://docs.sarvam.ai/api/api-guides-tutorials/dubbing/faq.md): Frequently asked questions about the Sarvam AI Dubbing API. Get answers about languages, voice cloning, export formats, watermarks, polling, and troubleshooting. - [Voice Cloning API](https://docs.sarvam.ai/api/api-guides-tutorials/voice-cloning/overview.md): Create a cloned voice from a reference clip, then generate speech with its voice ID. - [How to prepare reference audio](https://docs.sarvam.ai/api/api-guides-tutorials/voice-cloning/how-to/prepare-reference-audio.md): Reference clip guidelines for the Sarvam AI Voice Cloning API. What makes a good 10-15 second reference, technical specs, documented limits, and what to avoid. - [How to clone across languages](https://docs.sarvam.ai/api/api-guides-tutorials/voice-cloning/how-to/clone-across-languages.md): Cross-lingual voice cloning with the Sarvam AI Voice Cloning API. Clone a voice in one language and have it speak any of the 13 supported Indian languages. - [How to choose audio formats](https://docs.sarvam.ai/api/api-guides-tutorials/voice-cloning/how-to/choose-audio-formats.md): Input and output audio formats for the Sarvam AI Voice Cloning API. Supported reference codecs, output codecs, sample rates, and file-size guidance. - [How to use quality control](https://docs.sarvam.ai/api/api-guides-tutorials/voice-cloning/how-to/use-quality-control.md): The built-in quality control (QC) pipeline of the Sarvam AI Voice Cloning API: ASR verification, character error rate scoring, and prompt-leak detection. What QC checks, when to disable it, and what you'll see. - [Supported languages](https://docs.sarvam.ai/api/api-guides-tutorials/voice-cloning/supported-languages.md): All language codes accepted by the Sarvam AI Voice Cloning API as language_code, with native-script guidance and code-mixed text support. - [FAQs](https://docs.sarvam.ai/api/api-guides-tutorials/voice-cloning/faq.md): Frequently asked questions about the Sarvam AI Voice Cloning API. Get answers about reference clips, saved voices, quality control, limits, and troubleshooting. - [Build Your First Voice Agent using LiveKit](https://docs.sarvam.ai/api/integration/build-voice-agent-with-live-kit.md): A beginner-friendly guide to building a real-time voice agent using LiveKit and Sarvam AI. Support for 11 languages (10 Indian + English) with natural voices and multilingual conversations. - [Build Your First Voice Agent using Pipecat](https://docs.sarvam.ai/api/integration/build-voice-agent-with-pipecat.md): A beginner-friendly guide to building a real-time voice agent using Pipecat and Sarvam AI. Support for 11 languages (10 Indian + English) with natural voices and multilingual conversations. - [Build a Voice Agent or WhatsApp Bot using Twilio](https://docs.sarvam.ai/api/integration/build-voice-agent-with-twilio.md): A beginner-friendly guide to building a real-time phone voice agent or a WhatsApp bot using Twilio and Sarvam AI. Support for 11 languages (10 Indian + English) with natural voices and multilingual conversations. - [Build Your First Voice Agent using Exotel](https://docs.sarvam.ai/api/integration/build-voice-agent-with-exotel.md): A beginner-friendly guide to building a real-time phone voice agent using Exotel, Pipecat, and Sarvam AI. Support for 11 languages (10 Indian + English) with natural voices and multilingual conversations. - [Build a Voice Agent using Vobiz](https://docs.sarvam.ai/api/integration/build-voice-agent-with-vobiz.md): A beginner-friendly guide to building a real-time phone voice agent using Vobiz and Sarvam AI. Support for 11 languages (10 Indian + English) with natural voices and multilingual conversations. - [Build Your First Voice Agent using Vapi](https://docs.sarvam.ai/api/integration/build-voice-agent-with-vapi.md): A beginner-friendly guide to building a voice agent on Vapi powered end-to-end by Sarvam AI: Saaras v4 speech-to-text, Sarvam-105B, and Bulbul v3 text-to-speech. Support for 11 languages (10 Indian + English). - [Build workflows with Sarvam AI in n8n](https://docs.sarvam.ai/api/integration/n8n.md): A beginner-friendly guide to automating Indian-language speech and chat in n8n using the Sarvam AI community node. Install, credentials, sample workflows, and patterns that mirror our LiveKit and Pipecat integrations. - [Build with Sarvam AI in the Vercel AI SDK](https://docs.sarvam.ai/api/integration/vercel-ai-sdk.md): Use the community sarvam-ai-sdk provider to call Sarvam chat, text-to-speech, speech-to-text, translation, transliteration, and language identification from the Vercel AI SDK: install, setup, and code examples for every capability. - [Build with Sarvam AI in LangChain](https://docs.sarvam.ai/api/integration/langchain.md): Use the official langchain-sarvam package to call Sarvam chat models from LangChain: install, setup, and code examples for invoke, streaming, tool calling, structured output, and reasoning mode. - [LiveKit in Production](https://docs.sarvam.ai/api/integration/livekit-production-best-practices.md): A tuning and operations guide for teams shipping cascading Sarvam STT, LLM, and TTS voice agents on the LiveKit Agents SDK: VAD, turn detection, endpointing, barge-in, and latency tuning. - [Pipecat Production Guide](https://docs.sarvam.ai/api/integration/pipecat-production-guide.md): A reference and tuning playbook for the Sarvam STT, TTS, and LLM plugins in Pipecat: the defaults that hurt you in production, latency budgets, turn-taking, and reliability, verified against source. - [MCP Server](https://docs.sarvam.ai/api/developer-tools/mcp.md): Connect Claude Desktop, Claude Code, Cursor, Windsurf, Zed, and other AI clients to Sarvam through MCP. Call every Sarvam API as a tool. - [Markdown & llms.txt](https://docs.sarvam.ai/api/developer-tools/llms-txt.md): Use Sarvam's llms.txt, llms-full.txt, and per-page Markdown to feed accurate, up-to-date documentation into any LLM, and learn when to use llms.txt vs the MCP server. - [Context7](https://docs.sarvam.ai/api/developer-tools/context7.md): Pull up-to-date Sarvam AI documentation into Claude Code, Cursor, Windsurf, and other AI clients through Context7, the Sarvam docs are pre-indexed, no setup on our side required. - [Agent Skills](https://docs.sarvam.ai/api/developer-tools/agent-skills.md): Install ready-made Agent Skills for Sarvam's SDKs into Claude Code, Cursor, Windsurf, and other AI coding assistants so they generate correct Sarvam API calls on the first try. - [Call Analytics Pipeline](https://docs.sarvam.ai/api/cookbook/guides/call-analytics-pipeline.md): Build a production-ready call analytics pipeline with Sarvam AI: batch speech-to-text with speaker diarization, structured LLM analysis, follow-up Q&A, and automated summaries. - [Live Video Transcription](https://docs.sarvam.ai/api/cookbook/guides/live-video-transcription.md): Build a real-time transcription and translation demo with Sarvam AI's Streaming Speech-to-Text API, Flask-SocketIO, and the browser's Web Audio API. - [Podcast Generator](https://docs.sarvam.ai/api/cookbook/guides/podcast-generator.md): Build a full-stack app that converts PDF documents into natural two-host podcasts using Sarvam Document Digitization, a Sarvam chat model, and Bulbul v3 TTS, in 11 languages (10 Indian + English). - [WhatsApp Bot](https://docs.sarvam.ai/api/cookbook/guides/whats-app-bot.md): Build a WhatsApp bot that replies to text messages and voice notes using Sarvam AI's speech-to-text, chat, and text-to-speech APIs, in 11 languages (10 Indian + English). - [Collection Agent using LiveKit](https://docs.sarvam.ai/api/cookbook/example-voice-agents/collection-agent.md): Build a voice-based collection agent for payment reminders and follow-ups using LiveKit and Sarvam AI. Support for 11 languages (10 Indian + English) with natural voices. - [Government Scheme Awareness Agent using LiveKit](https://docs.sarvam.ai/api/cookbook/example-voice-agents/government-scheme-agent.md): Build a voice-based agent that helps citizens understand and apply for government schemes using LiveKit and Sarvam AI. Support for 11 languages (10 Indian + English). - [Tutor Agent using Pipecat](https://docs.sarvam.ai/api/cookbook/example-voice-agents/tutor-agent.md): Build a voice-based tutor agent that teaches students in multiple Indian languages using Pipecat and Sarvam AI. Perfect for EdTech applications. - [Loan Advisory Agent using Pipecat](https://docs.sarvam.ai/api/cookbook/example-voice-agents/loan-advisory-agent.md): Build a voice-based loan advisory agent that helps customers understand loan options using Pipecat and Sarvam AI. Support for 11 languages (10 Indian + English). - [Self-Hosted Deployments](https://docs.sarvam.ai/api/self-hosted/introduction.md): Deploy Sarvam's speech and document-intelligence models inside your own AWS account with Amazon SageMaker. Your audio and documents never leave your VPC. - [Managed vs Self-Hosted](https://docs.sarvam.ai/api/self-hosted/hosted-vs-self-hosted.md): Same Sarvam models, two delivery models. Compare the Managed API at api.sarvam.ai with a self-hosted SageMaker deployment in your own AWS account, and pick the right one. - [How it works](https://docs.sarvam.ai/api/self-hosted/architecture.md): The architecture of a Sarvam self-hosted deployment: AWS Marketplace entitlement, a SageMaker model package, and an inference endpoint running in your own VPC. Plus how to choose real-time, async, or batch. - [Get started on SageMaker](https://docs.sarvam.ai/api/self-hosted/sagemaker/get-started.md): Prerequisites, IAM setup, AWS Marketplace subscription, and service quotas for deploying Sarvam models on Amazon SageMaker. Everything you need before your first endpoint. - [Deploy Speech-to-Text (Saaras v3)](https://docs.sarvam.ai/api/self-hosted/sagemaker/deploy-saaras.md): Deploy the Saaras v3 speech-to-text model on Amazon SageMaker — real-time and streaming endpoints — from your AWS Marketplace subscription using boto3. - [Deploy Sarvam Vision](https://docs.sarvam.ai/api/self-hosted/sagemaker/deploy-vision.md): Deploy the Sarvam Vision document-intelligence model on Amazon SageMaker — real-time, async, and batch — from your AWS Marketplace subscription. OCR and parse PDFs and images across 23 languages. - [Deploy Text-to-Speech (Bulbul v3)](https://docs.sarvam.ai/api/self-hosted/sagemaker/deploy-bulbul.md): Deploy the Bulbul v3 text-to-speech model on Amazon SageMaker — real-time, server-side streaming (SSE), and bidirectional endpoints — from your AWS Marketplace subscription using boto3. - [Deploy with Terraform](https://docs.sarvam.ai/api/self-hosted/sagemaker/terraform.md): Deploy a Sarvam SageMaker endpoint with Terraform. One apply provisions the IAM execution role, the model from your Marketplace package, the endpoint config, and a live endpoint. - [API reference — Speech-to-Text](https://docs.sarvam.ai/api/self-hosted/sagemaker/api-saaras.md): The InvokeEndpoint request and response contract for a self-hosted Saaras v3 endpoint on SageMaker: parameters, the five output modes, response schema, the streaming protocol, and errors. - [API reference — Sarvam Vision](https://docs.sarvam.ai/api/self-hosted/sagemaker/api-vision.md): The InvokeEndpoint request and response contract for a self-hosted Sarvam Vision endpoint on SageMaker: content types, custom-attribute options, limits, the JSON response envelope, and error codes. - [API reference — Text-to-Speech](https://docs.sarvam.ai/api/self-hosted/sagemaker/api-bulbul.md): The InvokeEndpoint request and response contract for a self-hosted Bulbul v3 endpoint on SageMaker: parameters, the three inference modes (real-time, server-side streaming, bidirectional), voices, codecs, and errors. - [Error handling on SageMaker](https://docs.sarvam.ai/api/self-hosted/sagemaker/errors.md): How every self-hosted Sarvam endpoint reports errors on SageMaker: the 424 collapse through InvokeEndpoint, the shared error envelope, and the overload back-off rule for Saaras v3, Bulbul v3, and Sarvam Vision. - [Configure & tune](https://docs.sarvam.ai/api/self-hosted/sagemaker/configure.md): Tune a Sarvam SageMaker endpoint with container environment variables — concurrency, default output mode, and GPU behaviour — set at model-creation time. - [Operations](https://docs.sarvam.ai/api/self-hosted/sagemaker/operations.md): Run Sarvam SageMaker endpoints in production: instance sizing and pricing, autoscaling and scale-to-zero, CloudWatch monitoring, security and data residency, and updating model versions. - [Limitations](https://docs.sarvam.ai/api/self-hosted/sagemaker/limitations.md): Known constraints of Sarvam self-hosted SageMaker deployments: network isolation, payload and duration limits, page caps, and region-specific packages. Read before you build. - [Troubleshooting](https://docs.sarvam.ai/api/self-hosted/sagemaker/troubleshooting.md): Fix common issues with Sarvam SageMaker deployments: failed model creation, endpoints stuck creating, invoke errors, latency and throttling. Diagnose with CloudWatch. - [Migrating from ElevenLabs: Overview](https://docs.sarvam.ai/api/migrations/from-elevenlabs/overview.md): Common setup changes for teams moving from ElevenLabs to Sarvam: base URL, auth header, SDK client, and error handling, shared across every guide in this section. - [Migrating Text-to-Speech from ElevenLabs to Sarvam](https://docs.sarvam.ai/api/migrations/from-elevenlabs/text-to-speech.md): Endpoint, parameter, and response-format differences between ElevenLabs and Sarvam's Text-to-Speech API, with before/after code for REST, streaming, and pronunciation dictionaries. - [Migrating Speech-to-Text from ElevenLabs to Sarvam](https://docs.sarvam.ai/api/migrations/from-elevenlabs/speech-to-text.md): Endpoint, parameter, and response-format differences between ElevenLabs Scribe and Sarvam's Speech-to-Text API, with before/after code for REST, batch, and streaming transcription. - [Migrating Voice Cloning from ElevenLabs to Sarvam](https://docs.sarvam.ai/api/migrations/from-elevenlabs/voice-cloning.md): How to move cloned voices from ElevenLabs Instant Voice Cloning to Sarvam's Voice Cloning API: endpoint mapping, reference-clip guidance, before/after code, and the capabilities that have no ElevenLabs analog. - [Migrating from Cartesia: Overview](https://docs.sarvam.ai/api/migrations/from-cartesia/overview.md): Common setup changes for teams moving from Cartesia to Sarvam: base URL, auth header, SDK client, and error handling, shared across every guide in this section. - [Migrating Text-to-Speech from Cartesia to Sarvam](https://docs.sarvam.ai/api/migrations/from-cartesia/text-to-speech.md): Endpoint, parameter, and response-format differences between Cartesia's Sonic Text-to-Speech API and Sarvam's Text-to-Speech API, with before/after code for REST and streaming. - [Migrating Speech-to-Text from Cartesia to Sarvam](https://docs.sarvam.ai/api/migrations/from-cartesia/speech-to-text.md): Endpoint, parameter, and response-format differences between Cartesia's Ink Speech-to-Text API and Sarvam's Speech-to-Text API, with before/after code for REST, batch, and streaming transcription. - [Migrating from Deepgram: Overview](https://docs.sarvam.ai/api/migrations/from-deepgram/overview.md): Common setup changes for teams moving from Deepgram to Sarvam: base URL, auth header, SDK client, and error handling, shared across every guide in this section. - [Migrating Text-to-Speech from Deepgram to Sarvam](https://docs.sarvam.ai/api/migrations/from-deepgram/text-to-speech.md): Endpoint, parameter, and response-format differences between Deepgram's Aura Text-to-Speech API and Sarvam's Text-to-Speech API, with before/after code for REST and streaming. - [Migrating Speech-to-Text from Deepgram to Sarvam](https://docs.sarvam.ai/api/migrations/from-deepgram/speech-to-text.md): Endpoint, parameter, and response-format differences between Deepgram's Nova Speech-to-Text API and Sarvam's Speech-to-Text API, with before/after code for REST, batch, and streaming transcription. - [Migrating from Gemini: Overview](https://docs.sarvam.ai/api/migrations/from-gemini/overview.md): Common setup changes for teams moving from the Gemini API to Sarvam: base URL, auth header, SDK client, and error handling, shared across every guide in this section. - [Migrating Chat Completion from Gemini to Sarvam](https://docs.sarvam.ai/api/migrations/from-gemini/chat-completion.md): Endpoint, parameter, and response-format differences between the Gemini API and Sarvam's Chat Completion API, with before/after code for both the OpenAI-compatible endpoint and each platform's native SDK. - [Migrating Text-to-Speech from Gemini to Sarvam](https://docs.sarvam.ai/api/migrations/from-gemini/text-to-speech.md): Endpoint, parameter, and response-format differences between Gemini's native audio text-to-speech generation and Sarvam's Text-to-Speech API, with before/after code. - [Migrating Speech-to-Text from Gemini to Sarvam](https://docs.sarvam.ai/api/migrations/from-gemini/speech-to-text.md): Endpoint, parameter, and response-format differences between transcribing audio through the Gemini API's audio input and Sarvam's dedicated Speech-to-Text API, with before/after code.