> For clean Markdown of any page, append `.md` to the page URL. > For a complete documentation index, see https://docs.sarvam.ai/llms.txt. > For full documentation content in one file, see https://docs.sarvam.ai/llms-full.txt. > For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.sarvam.ai/_mcp/server. # Models > Complete overview of Sarvam AI's specialized models for Indian languages. Choose the right model for your use case - from speech processing to text generation, translation, and document intelligence. Sarvam AI provides a purpose-built AI stack for building applications in Indian languages. Our models span speech-to-text, speech translation, text translation, and high-quality text-to-speech, designed specifically for India's linguistic diversity, accents, and real-world usage patterns. Each model is trained and evaluated on Indian languages and culturally grounded data, enabling higher accuracy in production scenarios. With simple, well-documented APIs and predictable performance, developers can build, deploy, and scale India-first AI experiences without managing model complexity. > **Tip** > > New to building for Indian languages? Start with [Building for Indian Languages](/api/getting-started/building-for-india), a practical guide to language coverage, code-mixing, scripts, native numerals, 8kHz telephony audio, and pronunciation control. ## Model Selection Guide **Sarvam Models**: trained for Indian languages: | Model | API | Description | | ------------------------------------------------------------------------ | --------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | [Saaras v4](/api/getting-started/models/saaras) | Speech to Text | State-of-the-art ASR with 23 language support (22 Indian + Global/Indian English) and multiple output modes: transcribe, translate, verbatim, translit, codemix. Default, recommended model; v3 remains available. | | [Bulbul v3](/api/getting-started/models/bulbul) | Text to Speech | Natural-sounding voices for 11 languages (10 Indian + English) with customizable pitch, pace, and speaker options. | | [Sarvam Voice Cloning](/api/api-guides-tutorials/voice-cloning/overview) | Voice Cloning | Create a voice from a short reference clip and synthesize speech with it across 12 Indian languages - clone inline per request, or save the voice once and reuse its ID. | | [Sarvam-105B](/api/getting-started/models/sarvam-105b) | Chat Completion | 105B parameter flagship model, `sarvam-105b` for complex reasoning and agentic tasks, plus `sarvam-105b-conversations` for real-time dialogue and voice agents. | | [Mayura](/api/getting-started/models/mayura) | Text Translation | High-quality translation between 11 languages (10 Indian + English) with context preservation. | | [Sarvam-Translate](/api/getting-started/models/sarvam-translate) | Text Translation | Extended translation support for all 23 languages (22 Indian + English) with superior accuracy. | | [Sarvam Vision](/api/getting-started/models/sarvam-vision) | Document Intelligence | Extract and digitize content from documents in 23 languages with accurate OCR and structured output. | ## Open-Weight Models Open-weight models are available on **`/v2`** using the same Sarvam API key and credits. > **Note** > > GLM-5.3, Gemma 4 31B, and DeepSeek V4 Flash are **available in beta** and rolling out > gradually. See [Access to Beta APIs](/api-reference/beta-apis). | Model | Context window | Modality | Reasoning | | ----------------------------------------------------------------------------- | ---------------: | ------------------- | :-------: | | [GLM-5.3](/api/getting-started/models/openweight/glm-5-3) | 1,048,576 tokens | Text → text | Yes | | [Gemma 4 31B](/api/getting-started/models/openweight/gemma-4-31b) | 131,072 tokens | Text + image → text | No | | [DeepSeek V4 Flash](/api/getting-started/models/openweight/deepseek-v4-flash) | 1,048,576 tokens | Text → text | Yes | [Read more about the open-weight models →](/api/getting-started/models/openweight) ## Language Support Overview Language coverage varies by model. Check the table below before picking one. Full per-model tables are linked from each model's own page. | Model | Languages | Status | | --------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------- | ----------- | | **Saaras v4** (Speech to Text) | 23 (22 Indian + Global/Indian English), [full list](/api/getting-started/models/saaras#language-support) | Recommended | | **Sarvam Translate** (Text Translation) | 23 (22 Indian + English), [full list](/api/getting-started/models/sarvam-translate#language-support) | Active | | **Sarvam Vision** (Document Intelligence) | 23 (22 Indian + English), [full list](/api/getting-started/models/sarvam-vision#supported-languages) | Active | | **Bulbul v3** (Text to Speech) | 11 (10 Indian + English), [full list](/api/getting-started/models/bulbul#language-support) | Active | | **Sarvam Voice Cloning** | 12 Indian languages, [full list](/api/api-guides-tutorials/voice-cloning/supported-languages) | Active | | **Mayura** (Text Translation) | 11 (10 Indian + English), [full list](/api/getting-started/models/mayura#language-support) | Active | | **Sarvam-105B** (Chat LLM) | 11 (10 Indian + English), `sarvam-105b`, `sarvam-105b-conversations` | Active | | **GLM-5.3** (Chat LLM, [open-weight](/api/getting-started/models/openweight)) | See the model page for language coverage | Beta | | **Gemma 4 31B** (Chat LLM, [open-weight](/api/getting-started/models/openweight)) | See the model page for language coverage | Beta | | **DeepSeek V4 Flash** (Chat LLM, [open-weight](/api/getting-started/models/openweight)) | See the model page for language coverage | Beta | ### 23-language set (Saaras v4, Sarvam Translate, Sarvam Vision) | Language | Code | | Language | Code | | --------- | ------- | - | -------- | -------- | | Hindi | `hi-IN` | | Assamese | `as-IN` | | Bengali | `bn-IN` | | Urdu | `ur-IN` | | Kannada | `kn-IN` | | Nepali | `ne-IN` | | Malayalam | `ml-IN` | | Konkani | `kok-IN` | | Marathi | `mr-IN` | | Kashmiri | `ks-IN` | | Odia | `od-IN` | | Sindhi | `sd-IN` | | Punjabi | `pa-IN` | | Sanskrit | `sa-IN` | | Tamil | `ta-IN` | | Santali | `sat-IN` | | Telugu | `te-IN` | | Manipuri | `mni-IN` | | English | `en-IN` | | Bodo | `brx-IN` | | Gujarati | `gu-IN` | | Maithili | `mai-IN` | | | | | Dogri | `doi-IN` | ### 11-language set (Bulbul v3, Mayura, Sarvam-105B) | Language | Code | | Language | Code | | -------- | ------- | - | --------- | ------- | | Hindi | `hi-IN` | | Kannada | `kn-IN` | | Bengali | `bn-IN` | | Malayalam | `ml-IN` | | Tamil | `ta-IN` | | Marathi | `mr-IN` | | Telugu | `te-IN` | | Punjabi | `pa-IN` | | Gujarati | `gu-IN` | | Odia | `od-IN` | | English | `en-IN` | | | | --- ## Use Cases #### Voice Assistant ### Build a multilingual voice assistant 1. **Speech Input**: Use Saaras v4 with `mode="transcribe"` to convert user speech to text 2. **Understanding**: Process with Sarvam-105B for intelligent responses 3. **Speech Output**: Convert responses to speech with Bulbul Perfect for customer service, smart home devices, and accessibility applications. [Learn how to build a voice agent with LiveKit →](/api/integration/build-voice-agent-with-live-kit) #### Content Localization ### Localize content across Indian languages Build end-to-end multilingual experiences from audio to translated speech: 1. **Transcribe speech**: Use Saaras v4 with `mode="transcribe"` to accurately convert audio in Indian languages into text. 2. **Translate at scale**: Convert transcripts across 23 languages (22 Indian + English) with Sarvam Translate. 3. **Generate localized audio**: Turn translated text into natural, production-ready speech with Bulbul's high-quality voices. This gives you a complete pipeline, speech → text → translation → localized audio, enabling developers to deliver fully localized content for apps, videos, learning platforms, and product experiences with minimal engineering effort. #### Call Center Analytics ### Analyze multilingual customer interactions 1. **Speech Recognition**: Convert calls to text with Saaras v4 (`mode="transcribe"`) 2. **Translation**: Use Saaras v4 with `mode="translate"` for English output 3. **Insights**: Extract patterns and sentiment from conversations Essential for customer experience optimization and compliance monitoring. [Learn how to analyze calls with our analytics cookbook →](/api/cookbook/guides/call-analytics-pipeline) #### Educational Platform ### Create inclusive learning experiences 1. **Content Translation**: Make materials accessible in 23 languages (22 Indian + English) 2. **Audio Learning**: Generate pronunciation guides with Bulbul 3. **Interactive Chat**: Enable Q\&A with Sarvam-105B in native languages Perfect for online education, language learning, and skill development platforms. #### Document Processing ### Extract and digitize document content 1. **Document Upload**: Submit PDFs or images of documents in Indian languages 2. **Text Extraction**: Use Sarvam Vision for accurate OCR across 23 languages 3. **Structured Output**: Get clean HTML, Markdown, or JSON output with tables preserved Ideal for digitizing government forms, invoices, legal documents, and historical records in Indian languages. [Learn more about Document AI →](/api/api-guides-tutorials/document-intelligence/overview) #### Legacy Models The following models are still available but are being phased out. We recommend migrating to the newer models listed above. #### [Sarvam-M: Chat & Reasoning (Deprecated)](/api/getting-started/models/sarvam-m) 24B parameter multilingual chat model with hybrid reasoning. Deprecated and no longer available through the API. Migrate to Sarvam-105B. #### [Sarvam-30B: Chat LLM (Deprecated)](/api/getting-started/models/sarvam-30b) 30B parameter multilingual chat model with strong reasoning and conversational capabilities. Deprecated. Migrate to Sarvam-105B. > Complete overview of Sarvam AI's specialized models for Indian languages. Choose the right model for your use case - from speech processing to text generation, translation, and document intelligence. ## Docs - [Saaras](https://docs.sarvam.ai/api/getting-started/models/saaras.md): Saaras v3 and v4 - Domain-aware speech translation models that convert speech directly to English text with enhanced telephony support and intelligent entity preservation. - [Bulbul](https://docs.sarvam.ai/api/getting-started/models/bulbul.md): Bulbul v3 - High-quality multilingual text-to-speech model for Indian languages with natural prosody and 30+ speaker voices. - [Sarvam-105B](https://docs.sarvam.ai/api/getting-started/models/sarvam-105b.md): Sarvam-105B - 105B parameter flagship multilingual language model delivering state-of-the-art performance on Indian language understanding, reasoning, and generation tasks. - [Open-Weight Models](https://docs.sarvam.ai/api/getting-started/models/openweight.md): Open-weight models served on Sarvam infrastructure through the OpenAI-compatible chat completions API, using your existing Sarvam API key and credits. - [GLM-5.3](https://docs.sarvam.ai/api/getting-started/models/openweight/glm-5-3.md): GLM-5.3 is a large-scale reasoning model built for complex software engineering and long-horizon agent tasks. It supports text input and output with a 1,048,576-token context window. - [Gemma 4 31B](https://docs.sarvam.ai/api/getting-started/models/openweight/gemma-4-31b.md): Gemma 4 31B is a 31B-parameter instruction-tuned model for text and image understanding. It supports a 131,072-token context window and tool calling. - [DeepSeek V4 Flash](https://docs.sarvam.ai/api/getting-started/models/openweight/deepseek-v4-flash.md): DeepSeek V4 Flash is a general-purpose open-weight reasoning model with text input and output, a 1,048,576-token context window, and tool calling. - [Mayura](https://docs.sarvam.ai/api/getting-started/models/mayura.md): Mayura - Advanced multilingual translation model for Indian languages with customizable translation styles, script control, and intelligent code-mixed content handling. - [Sarvam Translate](https://docs.sarvam.ai/api/getting-started/models/sarvam-translate.md): Sarvam Translate - Comprehensive translation model supporting all 22 official Indian languages with formal translation style and structured text optimization. - [Sarvam Vision](https://docs.sarvam.ai/api/getting-started/models/sarvam-vision.md): Sarvam Vision - A 3B parameter multimodal model delivering world-class Document Intelligence and visual understanding with unmatched accuracy for 23 languages (22 Indian + English). - [Sarvam-30B (Deprecated)](https://docs.sarvam.ai/api/getting-started/models/sarvam-30b.md): Sarvam-30B - 30B parameter multilingual language model optimized for Indian languages with strong reasoning, coding, and conversational capabilities. Deprecated; migrate to Sarvam-105B. - [Sarvam-M (Deprecated)](https://docs.sarvam.ai/api/getting-started/models/sarvam-m.md): Sarvam-M - deprecated 24B parameter multilingual, hybrid-reasoning language model with strong Indian language benchmark performance. - [Saarika](https://docs.sarvam.ai/api/getting-started/models/saarika.md): Saarika v2.5 - High-accuracy speech recognition model for Indian languages with superior multi-speaker handling, telephony optimization, and automatic code-mixing support.