Skip to navigation

How to select output mode

View as Markdown

Saaras v3 and Saaras v4 support the same multiple output modes to handle different transcription and translation needs. Use the mode parameter to specify how you want the audio processed.

The mode parameter is available for Saaras v4 (default, recommended, latest) and Saaras v3. Mode behaviour is identical on both.

Output Mode Comparison

For the same input audio saying: “मेरा फोन नंबर है 9840950950” (My phone number is 9840950950)

ModeDescriptionExample Output
transcribeStandard transcription with number normalizationमेरा फोन नंबर है 9840950950
translateTranslate to EnglishMy phone number is 9840950950
verbatimExact word-for-word, preserves spoken numbersमेरा फोन नंबर है नौ आठ चार zero नौ पांच zero नौ पांच zero
translitRomanized/Latin scriptmera phone number hai 9840950950
codemixEnglish words in English, Indic words in native scriptमेरा phone number है 9840950950

When to Use Each Mode

ModeBest For
transcribeCall recordings, meetings, voice notes, general transcription
translateAnalytics dashboards, English-only systems, international teams
verbatimLegal transcriptions, compliance, preserving exact spoken content
translitSystems that only support Latin characters, search indexing
codemixHinglish conversations, mixed-language customer support

Example Code

Still available. All five modes behave as described above.

from sarvamai import SarvamAI
client = SarvamAI(api_subscription_key="YOUR_SARVAM_API_KEY")
# Standard transcription in original language
response = client.speech_to_text.transcribe(
file=open("audio.wav", "rb"),
model="saaras:v3",
language_code="hi-IN",
mode="transcribe" # Default mode
)
print(response.transcript)
# Output: मेरा फोन नंबर है 9840950950