> For clean Markdown of any page, append `.md` to the page URL. > For a complete documentation index, see https://docs.sarvam.ai/llms.txt. > For full documentation content in one file, see https://docs.sarvam.ai/llms-full.txt. > For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.sarvam.ai/_mcp/server. # Saaras > Saaras v3 and v4 - Domain-aware speech translation models that convert speech directly to English text with enhanced telephony support and intelligent entity preservation. Saaras is our state-of-the-art speech recognition family with flexible output formats. It supports multiple output modes including transcription, translation, verbatim, transliteration, and code-mixed outputs. Saaras is built to make Indic languages LLM-comprehensible, offering accurate transcriptions and translations across 23 languages (22 Indian languages + English). > **Note** > > **Saaras v4** is the latest model and is now the default, recommended model, adding Global English support (in addition to Indian English) across all output modes while keeping the same 22 Indic language coverage. **Saaras v3** remains available. Both are available on the **Speech-to-Text endpoint** (`/speech-to-text`) and support multiple output modes via the `mode` parameter. ## At a Glance | | | | --------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | **Model ID** | `saaras:v4` (default, recommended, latest), `saaras:v3` | | **What it does** | Speech-to-text with five output modes: transcribe, translate, verbatim, translit, codemix | | **Languages** | 23 (22 Indian + English), automatic language detection ([full list](#language-support)). `saaras:v4` additionally supports Global English alongside Indian English | | **APIs** | [REST](/api/api-guides-tutorials/speech-to-text/rest-api) (≤30 s), [Batch](/api/api-guides-tutorials/speech-to-text/batch-api) (≤2 hr/file), [WebSocket streaming](/api/api-guides-tutorials/speech-to-text/streaming-api) | | **Input limits** | 30 s per REST request; WAV, MP3, AAC, FLAC, OGG and more, [all limits](#limits) | | **Pricing** | [Pricing page](/api/getting-started/pricing) | | **Best for** | Voice agents, call analytics, 8 kHz telephony audio, code-mixed speech | | **Known limitations** | [See below](#known-limitations) | ## Saaras v3 vs v4 | | Saaras v3 | Saaras v4 | | ------------------- | -------------------------------------------------- | -------------------------------------------------- | | **Status** | | Default, recommended, latest | | **Output modes** | transcribe, translate, verbatim, translit, codemix | transcribe, translate, verbatim, translit, codemix | | **English support** | Indian English (`en-IN`) | Indian English **and** Global English | | **Endpoint** | `/speech-to-text` | `/speech-to-text` | > **Tip** > > `saaras:v4` is now the default model — you'll get it even if you omit the `model` parameter. To use the previous model instead, set `model="saaras:v3"` in any of the requests below. The request/response shape is identical across both. > **Tip** > > `saaras:v4` also supports **Keyterm Prompting**: a list of up to 50 domain-specific names, places, brands, or technical terms to bias recognition toward. The examples below pass `keyterms` — it is accepted only on `saaras:v4`, on the REST and Batch APIs, so drop that argument when using `saaras:v3` or the streaming endpoint. See [Keyterm Prompting](/api/api-guides-tutorials/speech-to-text/how-to/keyterms). ## Output Modes Saaras supports multiple output modes via the `mode` parameter, available on both `saaras:v3` and `saaras:v4`. Each mode produces different output formats for the same input audio. **Example audio:** *"मेरा फोन नंबर है 9840950950"* | Mode | Description | Example Output | | ---------------------- | --------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------- | | `transcribe` (default) | Standard transcription in the original language with proper formatting and number normalization | `मेरा फोन नंबर है 9840950950` | | `translate` | Translates speech from any supported Indic language to English | `My phone number is 9840950950` | | `verbatim` | Exact word-for-word transcription without normalization, preserving filler words and spoken numbers as-is | `मेरा फोन नंबर है नौ आठ चार zero नौ पांच zero नौ पांच zero` | | `translit` | Romanization - Transliterates speech to Latin/Roman script | `mera phone number hai 9840950950` | | `codemix` | Code-mixed text with English words in English and Indic words in native script | `मेरा phone number है 9840950950` | ## Key Features #### Domain-Aware Translation Advanced prompting system for domain-specific translation and hotword retention, ensuring accurate context preservation. #### Superior Telephony Performance Optimized for 8KHz telephony audio with enhanced multi-speaker recognition capabilities. #### Intelligent Entity Preservation Preserves proper nouns and entities accurately across languages, maintaining context and meaning. #### Multi-Language Support Supports 23 languages (22 Indian + English) with optional language identification. #### Speaker Diarization Provides diarized outputs with precise timestamps for multi-speaker conversations through batch API. #### Direct Translation Converts speech directly to English text, eliminating the need for separate transcription and translation steps. #### Global English (v4) `saaras:v4` extends English recognition beyond Indian English (`en-IN`) to Global English accents, on top of the same 22 Indic languages. ## Language Support Saaras v3 supports 23 languages (22 Indian languages + English) with comprehensive dialect and accent coverage, including code-mixed audio support and intelligent proper noun preservation for speech-to-English translation. | Language | Language Code | | Language | Language Code | | --------- | ------------- | - | -------- | ------------- | | Hindi | `hi-IN` | | Assamese | `as-IN` | | Bengali | `bn-IN` | | Urdu | `ur-IN` | | Kannada | `kn-IN` | | Nepali | `ne-IN` | | Malayalam | `ml-IN` | | Konkani | `kok-IN` | | Marathi | `mr-IN` | | Kashmiri | `ks-IN` | | Odia | `od-IN` | | Sindhi | `sd-IN` | | Punjabi | `pa-IN` | | Sanskrit | `sa-IN` | | Tamil | `ta-IN` | | Santali | `sat-IN` | | Telugu | `te-IN` | | Manipuri | `mni-IN` | | English | `en-IN` | | Bodo | `brx-IN` | | Gujarati | `gu-IN` | | Maithili | `mai-IN` | | | | | Dogri | `doi-IN` | > **Note** > > Language codes are optional. When not specified or set to `unknown`, the model will automatically detect the input language and return a `language_probability` score indicating detection confidence. **Additional Capabilities:** * Includes dialects and accents of the above languages * Code-mixed audio support * Intelligent Proper Noun and Entity Preservation to ensure proper nouns, regional names, and entities are recognized and retained accurately during transcription ## API Response Format The Speech-to-Text API returns a JSON response with the following fields: | Field | Type | Description | | ---------------------- | ---------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `request_id` | `string` | Unique identifier for the API request. | | `transcript` | `string` | The transcribed text from the provided audio file. | | `timestamps` | `object or null` | Contains chunk-level timestamps (`start_time_seconds`, `end_time_seconds`, `words`, each entry is a sentence/phrase, not an individual word). Only included when `with_timestamps` is set to `true`. Word-level timestamps are not supported. | | `diarized_transcript` | `object or null` | Diarized transcript with speaker labels. Available through batch API. | | `language_code` | `string or null` | BCP-47 code of the detected language (e.g., `hi-IN`). Returns the most predominant language if multiple are detected. Returns `null` if no language is detected. | | `language_probability` | `number or null` | Float value (0.0 to 1.0) indicating the probability of the detected language being correct. Higher values indicate higher confidence. Returns a value when `language_code` is not provided or set to `unknown`. Returns `null` when a specific `language_code` is provided (language detection is skipped). Always present in the response. | **Example Response:** ```json { "request_id": "20260209_abc123-def4-5678-ghij-klmnopqrstuv", "transcript": "नमस्ते, आप कैसे हैं?", "timestamps": null, "diarized_transcript": null, "language_code": "hi-IN", "language_probability": 0.95 } ``` ## Key Capabilities #### Transcribe Mode Standard transcription in the original language with proper formatting and number normalization. This is the default mode. #### Python ```python from sarvamai import SarvamAI client = SarvamAI( api_subscription_key="YOUR_SARVAM_API_KEY" ) try: response = client.speech_to_text.transcribe( file=open("audio.wav", "rb"), model="saaras:v4", mode="transcribe", # default mode keyterms=["Sarvam", "New Delhi", "Vistaar"] ) print(response.transcript) # Output: मेरा फोन नंबर है 9840950950 except Exception as e: print(f"Error: {e}") ``` #### JavaScript ```javascript import { SarvamAIClient } from "sarvamai"; import fs from "fs"; const API_KEY = "YOUR_SARVAM_API_KEY"; const FILE_PATH = "./audio.wav"; // Replace with your audio file path async function main() { const client = new SarvamAIClient({ apiSubscriptionKey: API_KEY }); try { const response = await client.speechToText.transcribe({ file: fs.createReadStream(FILE_PATH), model: "saaras:v4", mode: "transcribe", // default mode keyterms: ["Sarvam", "New Delhi", "Vistaar"] }); console.log(response.transcript); // Output: मेरा फोन नंबर है 9840950950 } catch (error) { console.error("Error:", error); } } main(); ``` #### cURL ```bash curl -X POST https://api.sarvam.ai/speech-to-text \ -H "api-subscription-key: " \ -H "Content-Type: multipart/form-data" \ -F file=@"audio.wav" \ -F model="saaras:v4" \ -F mode="transcribe" \ -F 'keyterms=["Sarvam","New Delhi","Vistaar"]' ``` #### Translate Mode Translates speech from any supported Indic language directly to English. Perfect for making Indic content LLM-comprehensible. #### Python ```python from sarvamai import SarvamAI client = SarvamAI( api_subscription_key="YOUR_SARVAM_API_KEY" ) try: response = client.speech_to_text.transcribe( file=open("audio.wav", "rb"), model="saaras:v4", mode="translate", keyterms=["Sarvam", "New Delhi", "Vistaar"] ) print(response.transcript) # Input: "मेरा फोन नंबर है 9840950950" # Output: My phone number is 9840950950 except Exception as e: print(f"Error: {e}") ``` #### JavaScript ```javascript import { SarvamAIClient } from "sarvamai"; import fs from "fs"; const API_KEY = "YOUR_SARVAM_API_KEY"; const FILE_PATH = "./audio.wav"; // Replace with your audio file path async function main() { const client = new SarvamAIClient({ apiSubscriptionKey: API_KEY }); try { const response = await client.speechToText.transcribe({ file: fs.createReadStream(FILE_PATH), model: "saaras:v4", mode: "translate", keyterms: ["Sarvam", "New Delhi", "Vistaar"] }); console.log(response.transcript); // Input: "मेरा फोन नंबर है 9840950950" // Output: My phone number is 9840950950 } catch (error) { console.error("Error:", error); } } main(); ``` #### cURL ```bash curl -X POST https://api.sarvam.ai/speech-to-text \ -H "api-subscription-key: " \ -H "Content-Type: multipart/form-data" \ -F file=@"audio.wav" \ -F model="saaras:v4" \ -F mode="translate" \ -F 'keyterms=["Sarvam","New Delhi","Vistaar"]' ``` #### Verbatim Mode Exact word-for-word transcription without normalization, preserving filler words and spoken numbers as-is. Ideal for detailed analysis. #### Python ```python from sarvamai import SarvamAI client = SarvamAI( api_subscription_key="YOUR_SARVAM_API_KEY" ) try: response = client.speech_to_text.transcribe( file=open("audio.wav", "rb"), model="saaras:v4", mode="verbatim", keyterms=["Sarvam", "New Delhi", "Vistaar"] ) print(response.transcript) # Input: "मेरा फोन नंबर है 9840950950" # Output: मेरा फोन नंबर है नौ आठ चार zero नौ पांच zero नौ पांच zero except Exception as e: print(f"Error: {e}") ``` #### JavaScript ```javascript import { SarvamAIClient } from "sarvamai"; import fs from "fs"; const API_KEY = "YOUR_SARVAM_API_KEY"; const FILE_PATH = "./audio.wav"; // Replace with your audio file path async function main() { const client = new SarvamAIClient({ apiSubscriptionKey: API_KEY }); try { const response = await client.speechToText.transcribe({ file: fs.createReadStream(FILE_PATH), model: "saaras:v4", mode: "verbatim", keyterms: ["Sarvam", "New Delhi", "Vistaar"] }); console.log(response.transcript); // Input: "मेरा फोन नंबर है 9840950950" // Output: मेरा फोन नंबर है नौ आठ चार zero नौ पांच zero नौ पांच zero } catch (error) { console.error("Error:", error); } } main(); ``` #### cURL ```bash curl -X POST https://api.sarvam.ai/speech-to-text \ -H "api-subscription-key: " \ -H "Content-Type: multipart/form-data" \ -F file=@"audio.wav" \ -F model="saaras:v4" \ -F mode="verbatim" \ -F 'keyterms=["Sarvam","New Delhi","Vistaar"]' ``` #### Translit Mode Romanization - Transliterates speech to Latin/Roman script. Useful for search indexing or when working with systems that don't support Indic scripts. #### Python ```python from sarvamai import SarvamAI client = SarvamAI( api_subscription_key="YOUR_SARVAM_API_KEY" ) try: response = client.speech_to_text.transcribe( file=open("audio.wav", "rb"), model="saaras:v4", mode="translit", keyterms=["Sarvam", "New Delhi", "Vistaar"] ) print(response.transcript) # Input: "मेरा फोन नंबर है 9840950950" # Output: mera phone number hai 9840950950 except Exception as e: print(f"Error: {e}") ``` #### JavaScript ```javascript import { SarvamAIClient } from "sarvamai"; import fs from "fs"; const API_KEY = "YOUR_SARVAM_API_KEY"; const FILE_PATH = "./audio.wav"; // Replace with your audio file path async function main() { const client = new SarvamAIClient({ apiSubscriptionKey: API_KEY }); try { const response = await client.speechToText.transcribe({ file: fs.createReadStream(FILE_PATH), model: "saaras:v4", mode: "translit", keyterms: ["Sarvam", "New Delhi", "Vistaar"] }); console.log(response.transcript); // Input: "मेरा फोन नंबर है 9840950950" // Output: mera phone number hai 9840950950 } catch (error) { console.error("Error:", error); } } main(); ``` #### cURL ```bash curl -X POST https://api.sarvam.ai/speech-to-text \ -H "api-subscription-key: " \ -H "Content-Type: multipart/form-data" \ -F file=@"audio.wav" \ -F model="saaras:v4" \ -F mode="translit" \ -F 'keyterms=["Sarvam","New Delhi","Vistaar"]' ``` #### Codemix Mode Code-mixed text with English words in English and Indic words in native script. Perfect for India's natural multilingual conversations. #### Python ```python from sarvamai import SarvamAI client = SarvamAI( api_subscription_key="YOUR_SARVAM_API_KEY" ) try: response = client.speech_to_text.transcribe( file=open("audio.wav", "rb"), model="saaras:v4", mode="codemix", keyterms=["Sarvam", "New Delhi", "Vistaar"] ) print(response.transcript) # Input: "मेरा फोन नंबर है 9840950950" # Output: मेरा phone number है 9840950950 except Exception as e: print(f"Error: {e}") ``` #### JavaScript ```javascript import { SarvamAIClient } from "sarvamai"; import fs from "fs"; const API_KEY = "YOUR_SARVAM_API_KEY"; const FILE_PATH = "./audio.wav"; // Replace with your audio file path async function main() { const client = new SarvamAIClient({ apiSubscriptionKey: API_KEY }); try { const response = await client.speechToText.transcribe({ file: fs.createReadStream(FILE_PATH), model: "saaras:v4", mode: "codemix", keyterms: ["Sarvam", "New Delhi", "Vistaar"] }); console.log(response.transcript); // Input: "मेरा फोन नंबर है 9840950950" // Output: मेरा phone number है 9840950950 } catch (error) { console.error("Error:", error); } } main(); ``` #### cURL ```bash curl -X POST https://api.sarvam.ai/speech-to-text \ -H "api-subscription-key: " \ -H "Content-Type: multipart/form-data" \ -F file=@"audio.wav" \ -F model="saaras:v4" \ -F mode="codemix" \ -F 'keyterms=["Sarvam","New Delhi","Vistaar"]' ``` ## Limits | Limit | Value | | ------------------------------------------------- | ------------------------------------------------------------------------------------------------ | | Max audio duration (real-time REST) | 30 seconds | | Supported formats | WAV, MP3, AAC, AIFF, OGG, OPUS, FLAC, MP4, AMR, WMA, WebM (auto-detected) | | Raw PCM input (`pcm_s16le`, `pcm_l16`, `pcm_raw`) | Requires `input_audio_codec`; must be 16 kHz | | Longer audio | Use the [Batch API](/api/api-guides-tutorials/speech-to-text/batch-api) (up to 2 hours per file) | | Rate limits | See [Rate Limits](/api/getting-started/ratelimits) | ## Known Limitations | Limitation | Detail | Workaround | | --------------------------------------- | -------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------- | | **30-second cap on real-time REST** | The real-time `/speech-to-text` endpoint only accepts audio up to 30 seconds long | Use the [Batch API](/api/api-guides-tutorials/speech-to-text/batch-api) for longer recordings (up to 2 hours per file) | | **`mode` parameter is v3/v4-only** | The `mode` parameter (transcribe / translate / verbatim / codemix) works only with `saaras:v3` and `saaras:v4` | Use `saaras:v3` or `saaras:v4`; older versions ignore `mode` | | **Translate mode outputs English only** | `mode="translate"` always produces English text, regardless of input language | Use transcribe mode if you need output in the original language | ## Next Steps #### [Developer quickstart](/api/api-guides-tutorials/speech-to-text/overview) Learn how to integrate Saaras v4 into your application. #### [API Reference](/api-reference/speech-to-text/transcribe) Complete API documentation for Speech-to-Text endpoint. #### [Cookbook](/api/api-guides-tutorials/speech-to-text/rest-api) Step-by-step tutorial for speech-to-text transcription. --- #### Saaras v2.5 (Deprecated Soon) > **Warning** > > **Deprecation Notice:** Saaras v2.5 will be deprecated soon. We recommend migrating to **Saaras v3** for improved accuracy and performance. The v2.5 model will continue to work during the transition period, but new features and improvements will only be available in v3. ### About Saaras v2.5 Saaras v2.5 is the previous speech translation model available in the **Speech-to-Text Translate endpoint** (`/speech-to-text-translate`). It converts speech directly to English text with enhanced telephony support and intelligent entity preservation. > **Note** > > **Key Difference:** Saaras v2.5 uses the `/speech-to-text-translate` endpoint, while Saaras v3 uses the `/speech-to-text` endpoint with mode parameter support. ### Key Features (v2.5) #### Domain-Aware Translation Advanced prompting system for domain-specific translation and hotword retention, ensuring accurate context preservation. #### Superior Telephony Performance Optimized for 8KHz telephony audio with enhanced multi-speaker recognition capabilities. #### Intelligent Entity Preservation Preserves proper nouns and entities accurately across languages, maintaining context and meaning. #### Multi-Language Support Supports 11 Indian languages with optional language identification. #### Speaker Diarization Provides diarized outputs with precise timestamps for multi-speaker conversations through batch API. #### Direct Translation Converts speech directly to English text, eliminating the need for separate transcription and translation steps. ### Translation Quality (v2.5 Benchmarks) COMET score, a robust metric for evaluating machine speech-translations, assesses semantic accuracy, fluency, and contextual relevance. Saaras v2.5 achieves exceptional performance on the Vistaar+Indicvoices Benchmark, a dataset curated from diverse Indian language audio sources, including code-mixed content, noisy environments, and regional accents. **COMET Score Performance:** * **Across 11 Languages:** 89.3% * **English:** 94.62% * **Hindi:** 91.83% * **9 Other languages:** 88.41% *Higher is better; Compared on VISTAAR + IndicVoices Benchmark* Why COMET? It evaluates not only lexical accuracy but also how well the translation captures meaning and context, critical for Indic languages with complex structures. **Dataset Description:** Contains real-world, multi-accented speech samples that covers 10 major Indic languages, ensuring representation of India's linguistic diversity. Includes code-mixed phrases, domain-specific vocabulary, and colloquial expressions. ### Migration Guide To migrate from Saaras v2.5 to v3: 1. **Change the endpoint:** Switch from `/speech-to-text-translate` to `/speech-to-text` 2. **Update the model parameter:** Change from `saaras:v2.5` to `saaras:v3` 3. **Add the mode parameter:** Use `mode="translate"` to get English output (similar to v2.5 behavior) ```diff # Endpoint change - POST /speech-to-text-translate + POST /speech-to-text # Parameter changes - model="saaras:v2.5" + model="saaras:v4" + mode="translate" ``` **SDK Migration:** ```diff # Python - response = client.speech_to_text.translate( - file=open("audio.wav", "rb"), model="saaras:v2.5" - ) + response = client.speech_to_text.transcribe( + file=open("audio.wav", "rb"), model="saaras:v4", mode="translate" + ) # JavaScript - const response = await client.speechToText.translate({ file: fs.createReadStream("audio.wav"), model: "saaras:v2.5" }); + const response = await client.speechToText.transcribe({ file: fs.createReadStream("audio.wav"), model: "saaras:v4", mode: "translate" }); ``` The response format remains compatible. > Saaras v3 and v4 - Domain-aware speech translation models that convert speech directly to English text with enhanced telephony support and intelligent entity preservation.