> For clean Markdown of any page, append `.md` to the page URL. > For a complete documentation index, see https://docs.sarvam.ai/llms.txt. > For full documentation content in one file, see https://docs.sarvam.ai/llms-full.txt. > For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.sarvam.ai/_mcp/server. # REST POST https://api.sarvam.ai/speech-to-text-translate Content-Type: multipart/form-data ## Speech to Text Translation API This API automatically detects the input language, transcribes the speech, and translates the text to English. ### Available Options: - **REST API** (Current Endpoint): For quick responses under 30 seconds with immediate results - **Batch API**: For longer audio files [Follow this documentation](https://docs.sarvam.ai/api-reference-docs/api-guides-tutorials/speech-to-text/batch-api) - Supports diarization (speaker identification) ### Note: - Pricing differs for REST and Batch APIs - Diarization is only available in Batch API with separate pricing - Please refer to [here](https://docs.sarvam.ai/api-reference-docs/pricing) for detailed pricing information Reference: https://docs.sarvam.ai/api-reference/legacy/speech-to-text-translate/translate ## Authentication - `api-subscription-key` header (required) — API Key authentication via header ## Request ### Body (multipart/form-data) This endpoint expects a multipart form containing a file. - `file` (file, required) — The audio file to transcribe. Supported formats include WAV, MP3, AAC, AIFF, OGG, OPUS, FLAC, MP4/M4A, AMR, WMA, WebM, and PCM formats. The API automatically detects most codec formats, but for PCM files (pcm_s16le, pcm_l16, pcm_raw), you must specify the input_audio_codec parameter. PCM files are supported only at 16kHz sample rate. Works best at 16kHz. Multiple channels will be merged. - `prompt` (string, optional) — Conversation context can be passed as a prompt to boost model accuracy. However, the current system is at an experimentation stage and doesn't match the prompt performance of large language models. - `model` (enum, optional) — Model to be used for speech to text translation. - **saaras:v2.5** (default): Translation model that translates audio from any spoken Indic language to English. - Example: Hindi audio → English text output For the latest model (saaras:v3), use the `/speech-to-text` endpoint with `mode="translate"`. - `input_audio_codec` (enum, optional) — Audio codec/format of the input file. Our API automatically detects all codec formats, but for PCM files specifically (pcm_s16le, pcm_l16, pcm_raw), you must pass this parameter. PCM files are supported only at 16kHz sample rate. ## Response ### 200 Successful Response - `request_id` (string, required, nullable) - `transcript` (string, required) — Transcript of the provided speech - `language_code` (enum, required, nullable) — This will return the BCP-47 code of language spoken in the input. If multiple languages are detected, this will return language code of most predominant spoken language. If no language is detected, this will be null - Allowed values: `hi-IN`, `bn-IN`, `kn-IN`, `ml-IN`, `mr-IN`, `od-IN`, `pa-IN`, `ta-IN`, `te-IN`, `gu-IN`, `en-IN`, `as-IN`, `ur-IN`, `ne-IN`, `kok-IN`, `ks-IN`, `sd-IN`, `sa-IN`, `sat-IN`, `mni-IN`, `brx-IN`, `mai-IN`, `doi-IN` - `diarized_transcript` (Sarvam_Model_API_DiarizedTranscript, optional, nullable) — Diarized transcript of the provided speech. Always `null` on the REST API — diarization is only available via the Batch API. - `language_probability` (double, optional, nullable) — Float value (0.0 to 1.0) indicating the probability of the detected language being correct. Higher values indicate higher confidence. **When it returns a value:** - When `language_code` is not provided in the request - When `language_code` is set to `unknown` **When it returns null:** - When a specific `language_code` is provided (language detection is skipped) The parameter is always present in the response. ## Errors ### 400 Bad Request Error Bad Request - `error` (Sarvam_Model_API_ErrorDetails, required) — Error details ### 403 Forbidden Error Forbidden - `error` (Sarvam_Model_API_ErrorDetails, required) — Error details ### 422 Unprocessable Entity Error Unprocessable Entity - `error` (Sarvam_Model_API_ErrorDetails, required) — Error details ### 429 Too Many Requests Error Quota Exceeded - `error` (Sarvam_Model_API_ErrorDetails, required) — Error details ### 500 Internal Server Error Internal Server Error - `error` (Sarvam_Model_API_ErrorDetails, required) — Error details ### 503 Service Unavailable Error Service Overloaded - `error` (Sarvam_Model_API_ErrorDetails, required) — Error details ## Types ### Sarvam_Model_API_DiarizedTranscript - `entries` (list of Sarvam_Model_API_DiarizedEntry, required) — List of diarized transcript entries. ### Sarvam_Model_API_ErrorDetails - `request_id` (string, required, nullable) - `message` (string, required) — Message describing the error - `code` (enum, required) — Error code for the specific error that has occurred. Refer to the error code documentation for more details. - Allowed values: `invalid_request_error`, `internal_server_error`, `unprocessable_entity_error`, `insufficient_quota_error`, `invalid_api_key_error`, `authentication_error`, `not_found_error`, `rate_limit_exceeded_error`, `model_call_error`, `gateway_timeout_error`, `billing_service_unavailable_error` ### Sarvam_Model_API_DiarizedEntry - `transcript` (string, required) — transcript of the segment of that audio - `start_time_seconds` (double, required) — Start time of the word in seconds. - `end_time_seconds` (double, required) — End time of the word in seconds. - `speaker_id` (string, required) — Speaker ID for the word. ## Examples **Request** ```json { "file": ">" } ``` **Response** ```json { "request_id": "20250101_0a1b2c3d-1234-5678-9abc-def012345678", "transcript": "Hello, how are you?", "language_code": "hi-IN", "diarized_transcript": null, "language_probability": 0.95 } ``` **SDK Code** ```typescript import { SarvamAIClient } from "sarvamai"; import fs from "fs"; const client = new SarvamAIClient({ apiSubscriptionKey: process.env.SARVAM_API_KEY, }); const response = await client.speechToText.translate({ file: fs.createReadStream("audio.wav"), }); ``` ```python import requests url = "https://api.sarvam.ai/speech-to-text-translate" files = { "file": "open('', 'rb')" } payload = { "input_audio_codec": , "model": , "prompt": } headers = {"api-subscription-key": ""} response = requests.post(url, data=payload, files=files, headers=headers) print(response.json()) ``` ```go package main import ( "fmt" "strings" "net/http" "io" ) func main() { url := "https://api.sarvam.ai/speech-to-text-translate" payload := strings.NewReader("-----011000010111000001101001\r\nContent-Disposition: form-data; name=\"file\"; filename=\"\"\r\nContent-Type: application/octet-stream\r\n\r\n\r\n-----011000010111000001101001\r\nContent-Disposition: form-data; name=\"input_audio_codec\"\r\n\r\n\r\n-----011000010111000001101001\r\nContent-Disposition: form-data; name=\"model\"\r\n\r\n\r\n-----011000010111000001101001\r\nContent-Disposition: form-data; name=\"prompt\"\r\n\r\n\r\n-----011000010111000001101001--\r\n") req, _ := http.NewRequest("POST", url, payload) req.Header.Add("api-subscription-key", "") res, _ := http.DefaultClient.Do(req) defer res.Body.Close() body, _ := io.ReadAll(res.Body) fmt.Println(res) fmt.Println(string(body)) } ``` ```ruby require 'uri' require 'net/http' url = URI("https://api.sarvam.ai/speech-to-text-translate") http = Net::HTTP.new(url.host, url.port) http.use_ssl = true request = Net::HTTP::Post.new(url) request["api-subscription-key"] = '' request.body = "-----011000010111000001101001\r\nContent-Disposition: form-data; name=\"file\"; filename=\"\"\r\nContent-Type: application/octet-stream\r\n\r\n\r\n-----011000010111000001101001\r\nContent-Disposition: form-data; name=\"input_audio_codec\"\r\n\r\n\r\n-----011000010111000001101001\r\nContent-Disposition: form-data; name=\"model\"\r\n\r\n\r\n-----011000010111000001101001\r\nContent-Disposition: form-data; name=\"prompt\"\r\n\r\n\r\n-----011000010111000001101001--\r\n" response = http.request(request) puts response.read_body ``` ```java import com.mashape.unirest.http.HttpResponse; import com.mashape.unirest.http.Unirest; HttpResponse response = Unirest.post("https://api.sarvam.ai/speech-to-text-translate") .header("api-subscription-key", "") .body("-----011000010111000001101001\r\nContent-Disposition: form-data; name=\"file\"; filename=\"\"\r\nContent-Type: application/octet-stream\r\n\r\n\r\n-----011000010111000001101001\r\nContent-Disposition: form-data; name=\"input_audio_codec\"\r\n\r\n\r\n-----011000010111000001101001\r\nContent-Disposition: form-data; name=\"model\"\r\n\r\n\r\n-----011000010111000001101001\r\nContent-Disposition: form-data; name=\"prompt\"\r\n\r\n\r\n-----011000010111000001101001--\r\n") .asString(); ``` ```php request('POST', 'https://api.sarvam.ai/speech-to-text-translate', [ 'multipart' => [ [ 'name' => 'file', 'filename' => '', 'contents' => null ] ] 'headers' => [ 'api-subscription-key' => '', ], ]); echo $response->getBody(); ``` ```csharp using RestSharp; var client = new RestClient("https://api.sarvam.ai/speech-to-text-translate"); var request = new RestRequest(Method.POST); request.AddHeader("api-subscription-key", ""); request.AddParameter("undefined", "-----011000010111000001101001\r\nContent-Disposition: form-data; name=\"file\"; filename=\"\"\r\nContent-Type: application/octet-stream\r\n\r\n\r\n-----011000010111000001101001\r\nContent-Disposition: form-data; name=\"input_audio_codec\"\r\n\r\n\r\n-----011000010111000001101001\r\nContent-Disposition: form-data; name=\"model\"\r\n\r\n\r\n-----011000010111000001101001\r\nContent-Disposition: form-data; name=\"prompt\"\r\n\r\n\r\n-----011000010111000001101001--\r\n", ParameterType.RequestBody); IRestResponse response = client.Execute(request); ``` ```swift import Foundation let headers = ["api-subscription-key": ""] let parameters = [ [ "name": "file", "fileName": "" ], [ "name": "input_audio_codec", "value": ], [ "name": "model", "value": ], [ "name": "prompt", "value": ] ] let boundary = "---011000010111000001101001" var body = "" var error: NSError? = nil for param in parameters { let paramName = param["name"]! body += "--\(boundary)\r\n" body += "Content-Disposition:form-data; name=\"\(paramName)\"" if let filename = param["fileName"] { let contentType = param["content-type"]! let fileContent = String(contentsOfFile: filename, encoding: String.Encoding.utf8) if (error != nil) { print(error as Any) } body += "; filename=\"\(filename)\"\r\n" body += "Content-Type: \(contentType)\r\n\r\n" body += fileContent } else if let paramValue = param["value"] { body += "\r\n\r\n\(paramValue)" } } let request = NSMutableURLRequest(url: NSURL(string: "https://api.sarvam.ai/speech-to-text-translate")! as URL, cachePolicy: .useProtocolCachePolicy, timeoutInterval: 10.0) request.httpMethod = "POST" request.allHTTPHeaderFields = headers request.httpBody = postData as Data let session = URLSession.shared let dataTask = session.dataTask(with: request as URLRequest, completionHandler: { (data, response, error) -> Void in if (error != nil) { print(error as Any) } else { let httpResponse = response as? HTTPURLResponse print(httpResponse) } }) dataTask.resume() ```