> For clean Markdown of any page, append `.md` to the page URL. > For a complete documentation index, see https://docs.sarvam.ai/llms.txt. > For full documentation content in one file, see https://docs.sarvam.ai/llms-full.txt. > For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.sarvam.ai/_mcp/server. # Saarika > Saarika v2.5 - High-accuracy speech recognition model for Indian languages with superior multi-speaker handling, telephony optimization, and automatic code-mixing support. Saarika-v2.5 is our flagship speech recognition model, specifically designed for Indian languages and accents. It always transcribes the audio in the same language it was spoken. It excels in handling complex multi-speaker conversations, telephony audio, and code-mixed speech with superior accuracy across 11 languages. > **Warning** > > **Deprecation Notice:** Saarika v2.5 will be deprecated soon. For transcription features, we recommend using [**Saaras v4**](/api/getting-started/models/saaras) with `mode="transcribe"`, which offers improved accuracy and additional output modes. ## Key Features #### Superior Telephony Performance Optimized for 8KHz telephony audio with enhanced noise handling and superior multi-speaker recognition capabilities. #### Intelligent Entity Preservation Preserves proper nouns and entities accurately across languages, maintaining context and meaning in transcriptions. #### Automatic Language Detection Optional automatic language identification with LID output. Use "unknown" when language is not known for automatic detection. #### Speaker Diarization Provides diarized outputs with precise timestamps for multi-speaker conversations through batch API processing. #### Automatic Code Mixing Intelligently handles mid-sentence language switches in code-mixed speech, perfect for India's multilingual conversations. #### Multi-Language Support Comprehensive support for Indian languages with high accuracy in mixed-language environments. ## Language Support Saarika supports 11 languages with comprehensive dialect and accent coverage, including code-mixed audio support and intelligent proper noun preservation. | Language | Language Code | | --------- | ------------- | | English | `en-IN` | | Hindi | `hi-IN` | | Bengali | `bn-IN` | | Tamil | `ta-IN` | | Telugu | `te-IN` | | Gujarati | `gu-IN` | | Kannada | `kn-IN` | | Malayalam | `ml-IN` | | Marathi | `mr-IN` | | Punjabi | `pa-IN` | | Odia | `od-IN` | > **Note** > > For automatic language detection, use `language_code="unknown"`. The model will automatically identify the spoken language and return it in the response. ## Performance Benchmarks Saarika delivers exceptional accuracy across all supported languages, as measured on the VISTAAR Benchmark. ### CER (Character Error Rate) Scores *Lower is better - Compared on VISTAAR Benchmark* * **Across 11 Languages: 4.96%** * **English: 4.45%** * **Hindi: 4.42%** * **9 Other languages: 5.07%** ### WER (Word Error Rate) Scores *Lower is better - Compared on VISTAAR Benchmark* * **Across 11 Languages: 18.32%** * **English: 8.26%** * **Hindi: 11.81%** * **9 Other languages: 20.15%** ### Detailed CER Performance by Language CER (Character Error Rate) measures the percentage of characters that are wrong in a transcription. Lower scores are better, with 0% being perfect. ## Key Capabilities > **Warning** > > **Deprecated Model:** Saarika v2.5 is a deprecated model and code examples have been removed to avoid new integrations against it. We recommend using **Saaras v3** (`model="saaras:v3"`) with the `mode` parameter for the best accuracy and features. See the [Saaras documentation](/api/getting-started/models/saaras) for up-to-date code examples covering basic transcription, code-mixed speech, and automatic language detection. Saarika supports the same core transcription capabilities now available in Saaras v3: * **Basic transcription** with a specified `language_code` for single-language content. * **Code-mixed speech** with automatic detection of language switches within a sentence. * **Automatic language detection** via `language_code="unknown"`. For implementation, use Saaras v3 with `mode="transcribe"`. See the [Saaras documentation](/api/getting-started/models/saaras). ## Limits | Limit | Value | | ------------------------------------------------- | ------------------------------------------------------------------------------------------------ | | Max audio duration (real-time REST) | 30 seconds | | Supported formats | WAV, MP3, AAC, AIFF, OGG, OPUS, FLAC, MP4, AMR, WMA, WebM (auto-detected) | | Raw PCM input (`pcm_s16le`, `pcm_l16`, `pcm_raw`) | Requires `input_audio_codec`; must be 16 kHz | | Longer audio | Use the [Batch API](/api/api-guides-tutorials/speech-to-text/batch-api) (up to 2 hours per file) | | Rate limits | See [Rate Limits](/api/getting-started/ratelimits) | ## Known Limitations | Limitation | Detail | Workaround | | ----------------------------------- | --------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------- | | **30-second cap on real-time REST** | The real-time `/speech-to-text` endpoint only accepts audio up to 30 seconds long | Use the [Batch API](/api/api-guides-tutorials/speech-to-text/batch-api) for longer recordings (up to 2 hours per file) | ## Next Steps #### [Developer quickstart](/api/api-guides-tutorials/speech-to-text/overview) Learn how to integrate the Saarika API within your application. #### [API Reference](/api-reference/speech-to-text/transcribe) Complete API documentation for speech to text endpoints. #### [Cookbook](/api/api-guides-tutorials/speech-to-text/rest-api) Step-by-step tutorial for speech-to-text transcription. > Saarika v2.5 - High-accuracy speech recognition model for Indian languages with superior multi-speaker handling, telephony optimization, and automatic code-mixing support.