Skip to navigation

How to specify language codes

View as Markdown

The language_code parameter tells the STT model which language to expect in the audio. Using the correct language code improves transcription accuracy.

Supported Languages (Saaras v3 and v4)

Saaras v3 and Saaras v4 support the same 22 Indian languages with BCP-47 format codes. saaras:v4 additionally recognizes Global English alongside Indian English (en-IN):

LanguageCodeLanguageCode
Hindihi-INAssameseas-IN
Bengalibn-INUrduur-IN
Kannadakn-INNepaline-IN
Malayalamml-INKonkanikok-IN
Marathimr-INKashmiriks-IN
Odiaod-INSindhisd-IN
Punjabipa-INSanskritsa-IN
Tamilta-INSantalisat-IN
Telugute-INManipurimni-IN
Englishen-INBodobrx-IN
Gujaratigu-INMaithilimai-IN
Dogridoi-IN

Automatic Language Detection

To enable automatic language detection, pass unknown as the language_code parameter. The model will detect the language from the audio.

Best Practice: Always specify the language code when you know the language of the audio. This improves accuracy and reduces processing time. Use unknown only when the language is truly unknown.

Example Code

Still available. 22 Indian languages plus Indian English (en-IN).

from sarvamai import SarvamAI
client = SarvamAI(api_subscription_key="YOUR_SARVAM_API_KEY")
# Specify language for better accuracy
response = client.speech_to_text.transcribe(
file=open("audio.wav", "rb"),
model="saaras:v3",
language_code="ta-IN", # Tamil
mode="transcribe"
)
print(response.transcript)