> For clean Markdown of any page, append `.md` to the page URL. > For a complete documentation index, see https://docs.sarvam.ai/llms.txt. > For full documentation content in one file, see https://docs.sarvam.ai/llms-full.txt. > For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.sarvam.ai/_mcp/server. # Which Speech-to-Text API to Use > Compare Sarvam's Speech-to-Text APIs, REST, WebSocket, and Batch (plus their Speech-to-Text-Translate variants), and pick the right one for your audio length, latency, and feature needs. Sarvam gives you four ways to run speech recognition on its models: the **REST**, **Realtime**, **Batch**, and legacy **WebSocket** APIs. They differ in how you send audio, how fast you get results, the maximum audio they accept, and which features (diarization, timestamps, partial transcripts) are available. Use this page to pick one before you start integrating. > **Note** > > Every transport also supports translating to **English** via `mode="translate"` (see the Tip below). A separate legacy **Speech-to-Text-Translate** counterpart exists on `saaras:v2.5`, the transport trade-offs below are identical, only the output language and underlying model change. See [Speech-to-Text-Translate](/api-reference/legacy/speech-to-text-translate/translate). ## Quick decision #### [REST](/api/api-guides-tutorials/speech-to-text/rest-api) A short clip (≤30s) and you want the transcript back in one call. #### [Realtime](/api/api-guides-tutorials/speech-to-text/realtime-streaming) Live microphone or call audio for a voice agent, need partial transcripts as the user speaks, and mid-call reconfiguration. #### [Batch](/api/api-guides-tutorials/speech-to-text/batch-api) Long recordings (up to 2 hours), with diarization and timestamps. #### [WebSocket (legacy)](/api/api-guides-tutorials/speech-to-text/streaming-api) Existing WebSocket integrations (now defaults to `saaras:v4`; `saaras:v3` still accepted), but superseded by Realtime for new work. ## Comparison | | REST | Realtime | Batch | WebSocket (Legacy) | | ----------------------------- | --------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------- | -------------------------------------------------------------------------- | ------------------------------------------------------------------- | | **Endpoint** | `POST /speech-to-text` | `GET /speech-to-text-realtime/ws` | `POST /speech-to-text/job/v1` (job flow) | `GET /speech-to-text/ws` | | **Model** | `saaras:v4` (default) | `saaras:v3-realtime` (default) | `saaras:v4` (default) | `saaras:v4` (default) | | **Processing** | Synchronous | Real-time streaming | Asynchronous (job) | Real-time streaming | | **Max audio length** | 30 seconds | Continuous (chunked) | 2 hours per file | Continuous (chunked) | | **Files per request** | 1 | 1 stream | Up to 20 per job | 1 stream | | **Interim / partial results** | N/A | Yes, `transcript.partial` throughout the utterance | N/A | No, final per utterance only | | **Results** | Final transcript in the response | Final transcript per utterance (VAD or manual turn end) | Final transcript, downloaded when the job completes | Final transcript per utterance (on VAD end-of-speech or `flush()`) | | **Latency** | One round-trip after upload | Lowest, partials arrive while audio is still streaming | Highest, minutes, depending on queue and duration | Low, but only a final per utterance | | **Mid-call reconfiguration** | N/A | Yes, `config.update`, no reconnect | N/A | No, reconnect required | | **VAD tuning** | N/A | 3 millisecond-valued parameters | N/A | 10+ frame-count parameters | | **Speaker diarization** | No | No | Yes | No | | **Timestamps** | Yes (chunk-level, via `with_timestamps`) | Optional segment-level (`return_timestamps`) | Yes (chunk-level) | No | | **Audio formats** | All [supported formats](/api/api-guides-tutorials/speech-to-text/overview) (auto-detected; PCM at 16 kHz) | Raw PCM only (`linear16`, `linear32`, `mulaw`, `alaw`) | All [supported formats](/api/api-guides-tutorials/speech-to-text/overview) | **WAV and raw PCM only** (`wav`, `pcm_s16le`, `pcm_l16`, `pcm_raw`) | | **Output modes** | `transcribe`, `translate`, `verbatim`, `translit`, `codemix` | Same 5 modes (final transcript only; partials are always plain transcription) | Same 5 modes | Same 5 modes | | **Best for** | Short clips, voice commands, quick tests | Voice agents, live captions with barge-in, call streaming | Meetings, interviews, call-center recordings, bulk pipelines | Existing WebSocket integrations | ## When to use each **REST: `POST /speech-to-text`** * The audio is already captured and short (≤30 seconds). * You want one request, one response, no connection to manage. * Examples: voice search, push-to-talk commands, transcribing a short voice note. * [REST API guide →](/api/api-guides-tutorials/speech-to-text/rest-api) **Realtime: `GET /speech-to-text-realtime/ws`** * Audio arrives continuously from a mic, browser, or telephony stream, for a voice agent or live captioning product. * You need true interim transcripts as the user speaks (fast barge-in), millisecond-based VAD tuning, and the ability to change language, mode, or VAD settings mid-call without reconnecting. * Note: raw PCM only, no diarization. * [Realtime Streaming guide →](/api/api-guides-tutorials/speech-to-text/realtime-streaming) **Batch: `POST /speech-to-text/job/v1`** * Recordings are long (up to 2 hours) or you have many files (up to 20 per job). * You need **speaker diarization** and **chunk-level timestamps** (e.g. subtitles, meeting minutes). * Latency isn't critical. You submit a job and download results when it finishes. * [Batch API guide →](/api/api-guides-tutorials/speech-to-text/batch-api) **WebSocket (Legacy): `GET /speech-to-text/ws`** * You have an existing integration on this WebSocket. It now defaults to `saaras:v4`; `saaras:v3` is still accepted for existing integrations. * For new voice-agent or live-transcription work, prefer Realtime Streaming above: it adds partial transcripts, simpler VAD tuning, and live reconfiguration. * Note: only **WAV / raw PCM** is accepted, and results are **final per utterance**: there are no interim partials. See [finalization semantics](/api/api-guides-tutorials/speech-to-text/streaming-api). * [Streaming API (Legacy) guide →](/api/api-guides-tutorials/speech-to-text/streaming-api) > **Tip** > > Need English output regardless of the spoken language? Use the same transport on the regular Speech-to-Text endpoints with `mode="translate"`: `saaras:v4` (default, recommended) or `v3` on REST, Batch, and the legacy WebSocket (`transcribe()`/`connect()`); `saaras:v3-realtime` (default) or `v4` on Realtime (the final-transcript `mode`). The dedicated **Speech-to-Text-Translate** endpoints are legacy (`saaras:v2.5`) and don't support `mode`. ## Related * [Speech-to-Text overview](/api/api-guides-tutorials/speech-to-text/overview) * [Supported audio formats & MIME types](/api/api-guides-tutorials/speech-to-text/overview) * [Credits & Rate Limits](/api/getting-started/ratelimits) > Compare Sarvam's Speech-to-Text APIs, REST, WebSocket, and Batch (plus their Speech-to-Text-Translate variants), and pick the right one for your audio length, latency, and feature needs.