> For clean Markdown of any page, append `.md` to the page URL. > For a complete documentation index, see https://docs.sarvam.ai/llms.txt. > For full documentation content in one file, see https://docs.sarvam.ai/llms-full.txt. > For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.sarvam.ai/_mcp/server. # Batch - Initiate Job POST https://api.sarvam.ai/speech-to-text-translate/job/v1 Content-Type: application/json Create a new speech to text translate bulk job and receive a job UUID and storage folder details for processing multiple audio files with translation. Set `job_parameters.input_audio_codec` when uploads are raw PCM (`pcm_s16le`, `pcm_l16`, or `pcm_raw`); the API auto-detects other formats. PCM must be 16 kHz. Reference: https://docs.sarvam.ai/api-reference/legacy/speech-to-text-translate/stt-translate/job/initiate ## Authentication - `api-subscription-key` header (required) — API Key authentication via header ## Request ### Query parameters - `ptu_id` (integer, optional, nullable) ### Headers - `api-subscription-key` (string, optional, default: ) — Your unique subscription key for authenticating requests to the Sarvam AI Speech-to-Text API. [Here are the steps to get your api key](https://docs.sarvam.ai/api-reference-docs/authentication#obtaining-your-api-subscription-key) ### Body (application/json) This endpoint expects a BulkJobInitRequestV1_SpeechToTextTranslateJobParameters_. - `job_parameters` (SpeechToTextTranslateJobParameters, required) — Job Parameters for the bulk job - `callback` (BulkJobCallback, optional, nullable) — Parameters for callback URL ## Response ### 202 Successful Response - `job_id` (string, required) — Job UUID. - `storage_container_type` (enum, required) — Storage Container Type - Allowed values: `Azure`, `Local`, `Google`, `Azure_V1` - `job_parameters` (BaseJobParameters, required) - `job_state` (enum, required) - Allowed values: `Accepted`, `Pending`, `Running`, `Completed`, `Failed` ## Errors ### 400 Bad Request Error Bad Request - `error` (ErrorDetails, required) — Error details ### 403 Forbidden Error Forbidden - `error` (ErrorDetails, required) — Error details ### 422 Unprocessable Entity Error Unprocessable Entity - `error` (ErrorDetails, required) — Error details ### 429 Too Many Requests Error Quota Exceeded - `error` (ErrorDetails, required) — Error details ### 500 Internal Server Error Internal Server Error - `error` (ErrorDetails, required) — Error details ### 503 Service Unavailable Error Service Overloaded - `error` (ErrorDetails, required) — Error details ## Types ### SpeechToTextTranslateJobParameters - `prompt` (string, optional, nullable) — Prompt to assist the transcription - `model` (enum, optional) — Model to be used for speech to text translation. - **saaras:v2.5** (default): Translation model that translates audio from any spoken Indic language to English. - Example: Hindi audio → English text output For the latest model (saaras:v3), use the `/speech-to-text` endpoint with `mode="translate"`. - Allowed values: `saaras:v2.5` - `with_diarization` (boolean, optional, default: false) — Enables speaker diarization, which identifies and separates different speakers in the audio. When set to true, the API will provide speaker-specific segments in the response. Note: This parameter is currently in Beta mode. - `num_speakers` (integer, optional, nullable) — Number of speakers to be detected in the audio. This is used when with_diarization is set to true. - `input_audio_codec` (enum, optional) — Audio codec/format of uploaded files. The API automatically detects most formats; for PCM files (pcm_s16le, pcm_l16, pcm_raw), you must specify this parameter. PCM files are supported only at 16kHz sample rate. - Allowed values: `wav`, `x-wav`, `wave`, `mp3`, `mpeg`, `mpeg3`, `x-mp3`, `x-mpeg-3`, `aac`, `x-aac`, `aiff`, `x-aiff`, `ogg`, `opus`, `flac`, `x-flac`, `mp4`, `x-m4a`, `amr`, `x-ms-wma`, `webm`, `pcm_s16le`, `pcm_l16`, `pcm_raw` ### BulkJobCallback - `url` (string, required) — Webhook url to call upon job completion - `auth_token` (string, optional, default: ) — Authorization token required for the callback Url ### BaseJobParameters ### ErrorDetails - `message` (string, required) — Message describing the error - `code` (enum, required) — Error code for the specific error that has occurred. Refer to the error code documentation for more details. - Allowed values: `invalid_request_error`, `internal_server_error`, `unprocessable_entity_error`, `insufficient_quota_error`, `invalid_api_key_error`, `authentication_error`, `rate_limit_exceeded_error`, `not_found_error` - `request_id` (string, optional, default: ) — Unique identifier for the request. Format: date_UUID4 ## Examples **Request** ```json {} ``` **Response** ```json { "job_id": "job_12345", "storage_container_type": "Azure_V1", "job_parameters": {}, "job_state": "Accepted" } ``` **SDK Code** ```typescript import { SarvamAIClient } from "sarvamai"; const client = new SarvamAIClient({ apiSubscriptionKey: process.env.SARVAM_API_KEY, }); const job = await client.speechToTextTranslateJob.initialise({ job_parameters: {}, }); ``` ```python import requests url = "https://api.sarvam.ai/speech-to-text-translate/job/v1" payload = {} headers = { "api-subscription-key": "", "Content-Type": "application/json" } response = requests.post(url, json=payload, headers=headers) print(response.json()) ``` ```go package main import ( "fmt" "strings" "net/http" "io" ) func main() { url := "https://api.sarvam.ai/speech-to-text-translate/job/v1" payload := strings.NewReader("{}") req, _ := http.NewRequest("POST", url, payload) req.Header.Add("api-subscription-key", "") req.Header.Add("Content-Type", "application/json") res, _ := http.DefaultClient.Do(req) defer res.Body.Close() body, _ := io.ReadAll(res.Body) fmt.Println(res) fmt.Println(string(body)) } ``` ```ruby require 'uri' require 'net/http' url = URI("https://api.sarvam.ai/speech-to-text-translate/job/v1") http = Net::HTTP.new(url.host, url.port) http.use_ssl = true request = Net::HTTP::Post.new(url) request["api-subscription-key"] = '' request["Content-Type"] = 'application/json' request.body = "{}" response = http.request(request) puts response.read_body ``` ```java import com.mashape.unirest.http.HttpResponse; import com.mashape.unirest.http.Unirest; HttpResponse response = Unirest.post("https://api.sarvam.ai/speech-to-text-translate/job/v1") .header("api-subscription-key", "") .header("Content-Type", "application/json") .body("{}") .asString(); ``` ```php request('POST', 'https://api.sarvam.ai/speech-to-text-translate/job/v1', [ 'body' => '{}', 'headers' => [ 'Content-Type' => 'application/json', 'api-subscription-key' => '', ], ]); echo $response->getBody(); ``` ```csharp using RestSharp; var client = new RestClient("https://api.sarvam.ai/speech-to-text-translate/job/v1"); var request = new RestRequest(Method.POST); request.AddHeader("api-subscription-key", ""); request.AddHeader("Content-Type", "application/json"); request.AddParameter("application/json", "{}", ParameterType.RequestBody); IRestResponse response = client.Execute(request); ``` ```swift import Foundation let headers = [ "api-subscription-key": "", "Content-Type": "application/json" ] let parameters = [] as [String : Any] let postData = JSONSerialization.data(withJSONObject: parameters, options: []) let request = NSMutableURLRequest(url: NSURL(string: "https://api.sarvam.ai/speech-to-text-translate/job/v1")! as URL, cachePolicy: .useProtocolCachePolicy, timeoutInterval: 10.0) request.httpMethod = "POST" request.allHTTPHeaderFields = headers request.httpBody = postData as Data let session = URLSession.shared let dataTask = session.dataTask(with: request as URLRequest, completionHandler: { (data, response, error) -> Void in if (error != nil) { print(error as Any) } else { let httpResponse = response as? HTTPURLResponse print(httpResponse) } }) dataTask.resume() ```