> For clean Markdown of any page, append `.md` to the page URL. > For a complete documentation index, see https://docs.sarvam.ai/llms.txt. > For full documentation content in one file, see https://docs.sarvam.ai/llms-full.txt. > For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.sarvam.ai/_mcp/server. # REST POST https://api.sarvam.ai/speech-to-text Content-Type: multipart/form-data ## Speech to Text API This API transcribes speech to text in multiple Indian languages and English. Supports transcription for interactive applications. ### Available Options: - **REST API** (Current Endpoint): For quick responses under 30 seconds with immediate results - **Batch API**: For longer audio files, [Follow This Documentation](https://docs.sarvam.ai/api-reference-docs/api-guides-tutorials/speech-to-text/batch-api) - Supports diarization (speaker identification) ### Note: - Pricing differs for REST and Batch APIs - Diarization is only available in Batch API with separate pricing - Please refer to [here](https://docs.sarvam.ai/api-reference-docs/pricing) for detailed pricing information Reference: https://docs.sarvam.ai/api-reference/speech-to-text/transcribe ## Authentication - `api-subscription-key` header (required) — API Key authentication via header ## Request ### Body (multipart/form-data) This endpoint expects a multipart form containing a file. - `file` (file, required) — The audio file to transcribe. Supported formats include WAV, MP3, AAC, AIFF, OGG, OPUS, FLAC, MP4/M4A, AMR, WMA, WebM, and PCM formats. The API automatically detects most codec formats, but for PCM files (pcm_s16le, pcm_l16, pcm_raw), you must specify the input_audio_codec parameter. PCM files are supported only at 16kHz sample rate. The API works best with audio files sampled at 16kHz. If the audio contains multiple channels, they will be merged into a single channel. - `model` (enum, optional) — Specifies the model to use for speech-to-text conversion. - **saaras:v4** (default, recommended, latest): Flexible output formats across all modes (transcribe, translate, verbatim, translit, codemix), supporting Global + Indian English and 22 Indic languages. - **saaras:v3**: State-of-the-art model with flexible output formats. Supports multiple modes via the `mode` parameter: transcribe, translate, verbatim, translit, codemix. - `mode` (enum, optional) — Mode of operation. **Only applicable when using saaras:v3 model.** Example audio: 'मेरा फोन नंबर है 9840950950' - **transcribe** (default): Standard transcription in the original language with proper formatting and number normalization. - Output: `मेरा फोन नंबर है 9840950950` - **translate**: Translates speech from any supported Indic language to English. - Output: `My phone number is 9840950950` - **verbatim**: Exact word-for-word transcription without normalization, preserving filler words and spoken numbers as-is. - Output: `मेरा फोन नंबर है नौ आठ चार zero नौ पांच zero नौ पांच zero` - **translit**: Romanization - Transliterates speech to Latin/Roman script only. - Output: `mera phone number hai 9840950950` - **codemix**: Code-mixed text with English words in English and Indic words in native script. - Output: `मेरा phone number है 9840950950` - `language_code` (enum, optional) — Specifies the language of the input audio in BCP-47 format. **Available Options:** - `unknown`: Use when the language is not known; the API will auto-detect. - `hi-IN`: Hindi - `bn-IN`: Bengali - `kn-IN`: Kannada - `ml-IN`: Malayalam - `mr-IN`: Marathi - `od-IN`: Odia - `pa-IN`: Punjabi - `ta-IN`: Tamil - `te-IN`: Telugu - `en-IN`: English - `gu-IN`: Gujarati - `as-IN`: Assamese - `ur-IN`: Urdu - `ne-IN`: Nepali - `kok-IN`: Konkani - `ks-IN`: Kashmiri - `sd-IN`: Sindhi - `sa-IN`: Sanskrit - `sat-IN`: Santali - `mni-IN`: Manipuri - `brx-IN`: Bodo - `mai-IN`: Maithili - `doi-IN`: Dogri - `with_timestamps` (boolean, optional) — Enables chunk-level timestamps in the response. If set to `true`, the response includes a `timestamps` object with `words`, `start_time_seconds`, and `end_time_seconds` (each entry covers a sentence or phrase, not an individual word). **Note:** Word-level timestamps are not supported. Speaker diarization is not supported on the REST API; use the Batch API for diarized, chunk-level timestamps. - `input_audio_codec` (enum, optional) — Input Audio codec/format of the input file. PCM files are supported only at 16kHz sample rate. - `keyterms` (list of string, optional) — List of up to 50 domain-specific terms (names, places, brands, technical terms) to bias recognition toward. Each keyterm can contain up to 64 characters. Put phrases such as `New Delhi` in one list item; do not send comma-separated terms in one string. Keyterms bias recognition — they do not guarantee that a term will appear in the transcript. **Only supported with `model=saaras:v4`.** Sent as a JSON-encoded array in one multipart form field, e.g. `keyterms=["Sarvam","New Delhi","Vistaar"]`. Do not use the older `keyterm` or `hotwords` fields. ## Response ### 200 Successful Response - `request_id` (string, required, nullable) - `transcript` (string, required) — The transcribed text from the provided audio file. - `language_code` (string, required, nullable) — This will return the BCP-47 code of language spoken in the input. If multiple languages are detected, this will return language code of most predominant spoken language. If no language is detected, this will be null - `timestamps` (Sarvam_Model_API_TimestampsModel, optional, nullable) — Chunk-level timestamps for the transcribed text (sentence/phrase segments, not individual words). Present only when `with_timestamps` is `true`; omitted otherwise. - `language_probability` (double, optional, nullable) — Float value (0.0 to 1.0) indicating confidence in the detected language. Present when `language_code` is omitted or set to `unknown`; omitted (or null) when a specific language code is provided. ## Errors ### 400 Bad Request Error Bad Request - `error` (Sarvam_Model_API_ErrorDetails, required) — Error details ### 403 Forbidden Error Forbidden - `error` (Sarvam_Model_API_ErrorDetails, required) — Error details ### 422 Unprocessable Entity Error Unprocessable Entity - `error` (Sarvam_Model_API_ErrorDetails, required) — Error details ### 429 Too Many Requests Error Quota Exceeded - `error` (Sarvam_Model_API_ErrorDetails, required) — Error details ### 500 Internal Server Error Internal Server Error - `error` (Sarvam_Model_API_ErrorDetails, required) — Error details ### 503 Service Unavailable Error Service Overloaded - `error` (Sarvam_Model_API_ErrorDetails, required) — Error details ## Types ### Sarvam_Model_API_TimestampsModel - `words` (list of string, required) — List of transcript chunks (sentence or phrase segments). Not individual words. - `start_time_seconds` (list of double, required) — List of start times for each chunk in seconds. - `end_time_seconds` (list of double, required) — List of end times for each chunk in seconds. ### Sarvam_Model_API_ErrorDetails - `request_id` (string, required, nullable) - `message` (string, required) — Message describing the error - `code` (enum, required) — Error code for the specific error that has occurred. Refer to the error code documentation for more details. - Allowed values: `invalid_request_error`, `internal_server_error`, `unprocessable_entity_error`, `insufficient_quota_error`, `invalid_api_key_error`, `authentication_error`, `not_found_error`, `rate_limit_exceeded_error`, `model_call_error`, `gateway_timeout_error`, `billing_service_unavailable_error` ## Examples **Request** ```json { "file": ">" } ``` **Response** ```json { "request_id": "20250101_0a1b2c3d-1234-5678-9abc-def012345678", "transcript": "नमस्ते, आप कैसे हैं?", "language_code": "hi-IN" } ``` **SDK Code** ```typescript import { SarvamAIClient } from "sarvamai"; import fs from "fs"; const client = new SarvamAIClient({ apiSubscriptionKey: process.env.SARVAM_API_KEY, }); const response = await client.speechToText.transcribe({ file: fs.createReadStream("audio.wav"), }); ``` ```swift import Foundation let url = URL(string: "https://api.sarvam.ai/speech-to-text")! let boundary = "Boundary-\(UUID().uuidString)" var request = URLRequest(url: url) request.httpMethod = "POST" request.setValue("YOUR_SARVAM_API_KEY", forHTTPHeaderField: "api-subscription-key") request.setValue("multipart/form-data; boundary=\(boundary)", forHTTPHeaderField: "Content-Type") let fileURL = URL(fileURLWithPath: "audio.wav") let fileData = try Data(contentsOf: fileURL) var body = Data() body.append("--\(boundary)\r\n".data(using: .utf8)!) body.append("Content-Disposition: form-data; name=\"file\"; filename=\"audio.wav\"\r\n".data(using: .utf8)!) body.append("Content-Type: audio/wav\r\n\r\n".data(using: .utf8)!) body.append(fileData) body.append("\r\n".data(using: .utf8)!) body.append("--\(boundary)--\r\n".data(using: .utf8)!) request.httpBody = body let task = URLSession.shared.dataTask(with: request) { data, response, error in if let error = error { print("Error:", error) return } if let data = data, let json = String(data: data, encoding: .utf8) { print(json) } } task.resume() ``` ```python import requests url = "https://api.sarvam.ai/speech-to-text" files = { "file": "open('', 'rb')" } payload = { "input_audio_codec": , "keyterms": , "language_code": , "mode": , "model": , "with_timestamps": } headers = {"api-subscription-key": ""} response = requests.post(url, data=payload, files=files, headers=headers) print(response.json()) ``` ```go package main import ( "fmt" "strings" "net/http" "io" ) func main() { url := "https://api.sarvam.ai/speech-to-text" payload := strings.NewReader("-----011000010111000001101001\r\nContent-Disposition: form-data; name=\"file\"; filename=\"\"\r\nContent-Type: application/octet-stream\r\n\r\n\r\n-----011000010111000001101001\r\nContent-Disposition: form-data; name=\"input_audio_codec\"\r\n\r\n\r\n-----011000010111000001101001\r\nContent-Disposition: form-data; name=\"keyterms\"\r\n\r\n\r\n-----011000010111000001101001\r\nContent-Disposition: form-data; name=\"language_code\"\r\n\r\n\r\n-----011000010111000001101001\r\nContent-Disposition: form-data; name=\"mode\"\r\n\r\n\r\n-----011000010111000001101001\r\nContent-Disposition: form-data; name=\"model\"\r\n\r\n\r\n-----011000010111000001101001\r\nContent-Disposition: form-data; name=\"with_timestamps\"\r\n\r\n\r\n-----011000010111000001101001--\r\n") req, _ := http.NewRequest("POST", url, payload) req.Header.Add("api-subscription-key", "") res, _ := http.DefaultClient.Do(req) defer res.Body.Close() body, _ := io.ReadAll(res.Body) fmt.Println(res) fmt.Println(string(body)) } ``` ```ruby require 'uri' require 'net/http' url = URI("https://api.sarvam.ai/speech-to-text") http = Net::HTTP.new(url.host, url.port) http.use_ssl = true request = Net::HTTP::Post.new(url) request["api-subscription-key"] = '' request.body = "-----011000010111000001101001\r\nContent-Disposition: form-data; name=\"file\"; filename=\"\"\r\nContent-Type: application/octet-stream\r\n\r\n\r\n-----011000010111000001101001\r\nContent-Disposition: form-data; name=\"input_audio_codec\"\r\n\r\n\r\n-----011000010111000001101001\r\nContent-Disposition: form-data; name=\"keyterms\"\r\n\r\n\r\n-----011000010111000001101001\r\nContent-Disposition: form-data; name=\"language_code\"\r\n\r\n\r\n-----011000010111000001101001\r\nContent-Disposition: form-data; name=\"mode\"\r\n\r\n\r\n-----011000010111000001101001\r\nContent-Disposition: form-data; name=\"model\"\r\n\r\n\r\n-----011000010111000001101001\r\nContent-Disposition: form-data; name=\"with_timestamps\"\r\n\r\n\r\n-----011000010111000001101001--\r\n" response = http.request(request) puts response.read_body ``` ```java import com.mashape.unirest.http.HttpResponse; import com.mashape.unirest.http.Unirest; HttpResponse response = Unirest.post("https://api.sarvam.ai/speech-to-text") .header("api-subscription-key", "") .body("-----011000010111000001101001\r\nContent-Disposition: form-data; name=\"file\"; filename=\"\"\r\nContent-Type: application/octet-stream\r\n\r\n\r\n-----011000010111000001101001\r\nContent-Disposition: form-data; name=\"input_audio_codec\"\r\n\r\n\r\n-----011000010111000001101001\r\nContent-Disposition: form-data; name=\"keyterms\"\r\n\r\n\r\n-----011000010111000001101001\r\nContent-Disposition: form-data; name=\"language_code\"\r\n\r\n\r\n-----011000010111000001101001\r\nContent-Disposition: form-data; name=\"mode\"\r\n\r\n\r\n-----011000010111000001101001\r\nContent-Disposition: form-data; name=\"model\"\r\n\r\n\r\n-----011000010111000001101001\r\nContent-Disposition: form-data; name=\"with_timestamps\"\r\n\r\n\r\n-----011000010111000001101001--\r\n") .asString(); ``` ```php request('POST', 'https://api.sarvam.ai/speech-to-text', [ 'multipart' => [ [ 'name' => 'file', 'filename' => '', 'contents' => null ] ] 'headers' => [ 'api-subscription-key' => '', ], ]); echo $response->getBody(); ``` ```csharp using RestSharp; var client = new RestClient("https://api.sarvam.ai/speech-to-text"); var request = new RestRequest(Method.POST); request.AddHeader("api-subscription-key", ""); request.AddParameter("undefined", "-----011000010111000001101001\r\nContent-Disposition: form-data; name=\"file\"; filename=\"\"\r\nContent-Type: application/octet-stream\r\n\r\n\r\n-----011000010111000001101001\r\nContent-Disposition: form-data; name=\"input_audio_codec\"\r\n\r\n\r\n-----011000010111000001101001\r\nContent-Disposition: form-data; name=\"keyterms\"\r\n\r\n\r\n-----011000010111000001101001\r\nContent-Disposition: form-data; name=\"language_code\"\r\n\r\n\r\n-----011000010111000001101001\r\nContent-Disposition: form-data; name=\"mode\"\r\n\r\n\r\n-----011000010111000001101001\r\nContent-Disposition: form-data; name=\"model\"\r\n\r\n\r\n-----011000010111000001101001\r\nContent-Disposition: form-data; name=\"with_timestamps\"\r\n\r\n\r\n-----011000010111000001101001--\r\n", ParameterType.RequestBody); IRestResponse response = client.Execute(request); ```