> For clean Markdown of any page, append `.md` to the page URL. > For a complete documentation index, see https://docs.sarvam.ai/llms.txt. > For full documentation content in one file, see https://docs.sarvam.ai/llms-full.txt. > For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.sarvam.ai/_mcp/server. # Batch - Initiate Job POST https://api.sarvam.ai/speech-to-text/job/v1 Content-Type: application/json Create a new speech to text bulk job and receive a job UUID and storage folder details for processing multiple audio files. Set `job_parameters.input_audio_codec` when uploads are raw PCM (`pcm_s16le`, `pcm_l16`, or `pcm_raw`); the API auto-detects other formats. PCM must be 16 kHz. Reference: https://docs.sarvam.ai/api-reference/speech-to-text/stt/job/initiate ## Authentication - `api-subscription-key` header (required) — API Key authentication via header ## Request ### Body (application/json) This endpoint expects a BulkJobInitRequestV1_SpeechToTextJobParameters_. - `job_parameters` (SpeechToTextJobParameters, required) — Job Parameters for the bulk job - `callback` (BulkJobCallback, optional, nullable) — Parameters for callback URL ## Response ### 202 Successful Response - `job_id` (string, required) — Job UUID. - `storage_container_type` (enum, required) — Storage Container Type - Allowed values: `Azure`, `Local`, `Google`, `Azure_V1` - `job_parameters` (BaseJobParameters, required) - `job_state` (enum, required) - Allowed values: `Accepted`, `Pending`, `Running`, `Completed`, `Failed` ## Errors ### 400 Bad Request Error Bad Request - `error` (ErrorDetails, required) — Error details ### 403 Forbidden Error Forbidden - `error` (ErrorDetails, required) — Error details ### 422 Unprocessable Entity Error Unprocessable Entity - `error` (ErrorDetails, required) — Error details ### 429 Too Many Requests Error Quota Exceeded - `error` (ErrorDetails, required) — Error details ### 500 Internal Server Error Internal Server Error - `error` (ErrorDetails, required) — Error details ### 503 Service Unavailable Error Service Overloaded - `error` (ErrorDetails, required) — Error details ## Types ### SpeechToTextJobParameters - `language_code` (enum, optional, nullable, default: unknown) — Specifies the language of the input audio in BCP-47 format. **Available Options:** - `unknown` (default): Use when the language is not known; the API will auto-detect. - `hi-IN`: Hindi - `bn-IN`: Bengali - `kn-IN`: Kannada - `ml-IN`: Malayalam - `mr-IN`: Marathi - `od-IN`: Odia - `pa-IN`: Punjabi - `ta-IN`: Tamil - `te-IN`: Telugu - `en-IN`: English - `gu-IN`: Gujarati **Additional Options (saaras:v3 only):** - `as-IN`: Assamese - `ur-IN`: Urdu - `ne-IN`: Nepali - `kok-IN`: Konkani - `ks-IN`: Kashmiri - `sd-IN`: Sindhi - `sa-IN`: Sanskrit - `sat-IN`: Santali - `mni-IN`: Manipuri - `brx-IN`: Bodo - `mai-IN`: Maithili - `doi-IN`: Dogri - Allowed values: `unknown`, `hi-IN`, `bn-IN`, `kn-IN`, `ml-IN`, `mr-IN`, `od-IN`, `pa-IN`, `ta-IN`, `te-IN`, `en-IN`, `gu-IN`, `as-IN`, `ur-IN`, `ne-IN`, `kok-IN`, `ks-IN`, `sd-IN`, `sa-IN`, `sat-IN`, `mni-IN`, `brx-IN`, `mai-IN`, `doi-IN` - `model` (enum, optional) — Model to be used for speech to text. - **saaras:v4** (default, recommended, latest): Flexible output formats across all modes (transcribe, translate, verbatim, translit, codemix), supporting Global + Indian English and 22 Indic languages. - **saaras:v3**: State-of-the-art model with flexible output formats. Supports multiple modes via the `mode` parameter: transcribe, translate, verbatim, translit, codemix. - Allowed values: `saaras:v3`, `saaras:v4` - `mode` (enum, optional, nullable, default: transcribe) — Mode of operation. **Only applicable when using saaras:v3 or saaras:v4 models.** Example audio: 'मेरा फोन नंबर है 9840950950' - **transcribe** (default): Standard transcription in the original language with proper formatting and number normalization. - Output: `मेरा फोन नंबर है 9840950950` - **translate**: Translates speech from any supported Indic language to English. - Output: `My phone number is 9840950950` - **verbatim**: Exact word-for-word transcription without normalization, preserving filler words and spoken numbers as-is. - Output: `मेरा फोन नंबर है नौ आठ चार zero नौ पांच zero नौ पांच zero` - **translit**: Romanization - Transliterates speech to Latin/Roman script only. - Output: `mera phone number hai 9840950950` - **codemix**: Code-mixed text with English words in English and Indic words in native script. - Output: `मेरा phone number है 9840950950` - Allowed values: `transcribe`, `translate`, `verbatim`, `translit`, `codemix` - `with_timestamps` (boolean, optional, default: false) — Whether to include timestamps in the response - `with_diarization` (boolean, optional, default: false) — Enables speaker diarization, which identifies and separates different speakers in the audio. In beta mode. - `num_speakers` (integer, optional, nullable) — Number of speakers to be detected in the audio. This is used when with_diarization is true. - `input_audio_codec` (enum, optional) — Audio codec/format of uploaded files. The API automatically detects most formats; for PCM files (pcm_s16le, pcm_l16, pcm_raw), you must specify this parameter. PCM files are supported only at 16kHz sample rate. - Allowed values: `wav`, `x-wav`, `wave`, `mp3`, `mpeg`, `mpeg3`, `x-mp3`, `x-mpeg-3`, `aac`, `x-aac`, `aiff`, `x-aiff`, `ogg`, `opus`, `flac`, `x-flac`, `mp4`, `x-m4a`, `amr`, `x-ms-wma`, `webm`, `pcm_s16le`, `pcm_l16`, `pcm_raw` - `keyterms` (list of string, optional, nullable) — List of up to 50 domain-specific terms (names, places, brands, technical terms) to bias recognition toward. Each keyterm can contain up to 64 characters. Put phrases such as `New Delhi` in one list item; do not send comma-separated terms in one string. Keyterms bias recognition — they do not guarantee that a term will appear in the transcript. **Only supported with `model=saaras:v4`.** Applied while transcribing every audio chunk in the job. Do not use the older `keyterm` or `hotwords` fields. ### BulkJobCallback - `url` (string, required) — Webhook url to call upon job completion - `auth_token` (string, optional, default: ) — Authorization token required for the callback Url ### BaseJobParameters ### ErrorDetails - `message` (string, required) — Message describing the error - `code` (enum, required) — Error code for the specific error that has occurred. Refer to the error code documentation for more details. - Allowed values: `invalid_request_error`, `internal_server_error`, `unprocessable_entity_error`, `insufficient_quota_error`, `invalid_api_key_error`, `authentication_error`, `rate_limit_exceeded_error`, `not_found_error` - `request_id` (string, optional, default: ) — Unique identifier for the request. Format: date_UUID4 ## Examples **Request** ```json {} ``` **Response** ```json { "job_id": "job_12345", "storage_container_type": "Azure_V1", "job_parameters": {}, "job_state": "Accepted" } ``` **SDK Code** ```typescript import { SarvamAIClient } from "sarvamai"; const client = new SarvamAIClient({ apiSubscriptionKey: process.env.SARVAM_API_KEY, }); const job = await client.speechToTextJob.initialise({ job_parameters: {}, }); ``` ```python import requests url = "https://api.sarvam.ai/speech-to-text/job/v1" payload = {} headers = { "api-subscription-key": "", "Content-Type": "application/json" } response = requests.post(url, json=payload, headers=headers) print(response.json()) ``` ```go package main import ( "fmt" "strings" "net/http" "io" ) func main() { url := "https://api.sarvam.ai/speech-to-text/job/v1" payload := strings.NewReader("{}") req, _ := http.NewRequest("POST", url, payload) req.Header.Add("api-subscription-key", "") req.Header.Add("Content-Type", "application/json") res, _ := http.DefaultClient.Do(req) defer res.Body.Close() body, _ := io.ReadAll(res.Body) fmt.Println(res) fmt.Println(string(body)) } ``` ```ruby require 'uri' require 'net/http' url = URI("https://api.sarvam.ai/speech-to-text/job/v1") http = Net::HTTP.new(url.host, url.port) http.use_ssl = true request = Net::HTTP::Post.new(url) request["api-subscription-key"] = '' request["Content-Type"] = 'application/json' request.body = "{}" response = http.request(request) puts response.read_body ``` ```java import com.mashape.unirest.http.HttpResponse; import com.mashape.unirest.http.Unirest; HttpResponse response = Unirest.post("https://api.sarvam.ai/speech-to-text/job/v1") .header("api-subscription-key", "") .header("Content-Type", "application/json") .body("{}") .asString(); ``` ```php request('POST', 'https://api.sarvam.ai/speech-to-text/job/v1', [ 'body' => '{}', 'headers' => [ 'Content-Type' => 'application/json', 'api-subscription-key' => '', ], ]); echo $response->getBody(); ``` ```csharp using RestSharp; var client = new RestClient("https://api.sarvam.ai/speech-to-text/job/v1"); var request = new RestRequest(Method.POST); request.AddHeader("api-subscription-key", ""); request.AddHeader("Content-Type", "application/json"); request.AddParameter("application/json", "{}", ParameterType.RequestBody); IRestResponse response = client.Execute(request); ``` ```swift import Foundation let headers = [ "api-subscription-key": "", "Content-Type": "application/json" ] let parameters = [] as [String : Any] let postData = JSONSerialization.data(withJSONObject: parameters, options: []) let request = NSMutableURLRequest(url: NSURL(string: "https://api.sarvam.ai/speech-to-text/job/v1")! as URL, cachePolicy: .useProtocolCachePolicy, timeoutInterval: 10.0) request.httpMethod = "POST" request.allHTTPHeaderFields = headers request.httpBody = postData as Data let session = URLSession.shared let dataTask = session.dataTask(with: request as URLRequest, completionHandler: { (data, response, error) -> Void in if (error != nil) { print(error as Any) } else { let httpResponse = response as? HTTPURLResponse print(httpResponse) } }) dataTask.resume() ```