Skip to navigation

Best Practice Guide for Bulbul v4 Flash

View as Markdown

Bulbul v4 Flash is Sarvam’s low-latency TTS tier for interactive workloads: voice agents, IVR, and anything where a listener is waiting for the first word. This guide applies only to model="bulbul:v4-flash" with persona speaker IDs (voice_language_style). For the stable bulbul:v3 model (short names like shubh, priya), see Bulbul v3 best practices.

Preview voices before you ship: Voices lists every v4 Flash persona with audio samples.


What is Bulbul v4 Flash?

Flash is the interactive tier of the v4 family. It runs on production infrastructure and is available over three transports (HTTP streaming, single-response REST, and WebSocket) so you can match your pipeline instead of forcing one mode.

CapabilityDetail
LanguagesHindi, Bengali, Gujarati, Kannada, Malayalam, Marathi, Odia, Punjabi, Tamil, Telugu, and English (Indian accent) — 11 codes, en-US is not accepted; use en-IN.
Code-mixingHinglish, Tanglish, and similar mixes when each language is written in its own script.
Indian namesStrong on Indian names, places, and cultural terms that global TTS often misreads.
Preprocessingenable_preprocessing normalises numbers, dates, currency, and mixed-script text before synthesis.
ProsodyIndependent pace, pitch, and loudness (v3 offers pace and temperature only).
CodecsWAV, MP3, FLAC, AAC; configurable MP3 bitrate on HTTP streaming.
Telephony8 kHz output for IVR and contact-centre legs.
Voice roster83 named personas, 224 speaker IDs on the API today.

Choosing the right endpoint

All transports serve the same model; they differ in latency, max sample rate, and how audio is delivered.

HTTP streamingSingle response (REST)WebSocket
PathPOST /text-to-speech/streamPOST /text-to-speechwss://api.sarvam.ai/text-to-speech/ws
ReturnsChunked audio bytesJSON with base64 in audios[0]{"type":"audio"} messages with base64 chunks
Max sample rate24 kHz48 kHz24 kHz
Playable before completionYesNoYes
Best forVoice agents, IVR, LLM read-aloudOffline files, broadcast-grade 48 kHzHigh-volume sequential synthesis on one connection

If a person is waiting to hear the first word, use HTTP streaming or WebSocket. The single-response endpoint finishes the full clip before replying, which listeners experience as dead air. Reserve REST for offline generation and when you need 32 / 44.1 / 48 kHz.

Reuse connections: Each new HTTPS connection pays DNS, TCP, and TLS before any audio work. Enable keep-alive or pool clients so one connection serves many synthesis calls.

See also Which API to Use.


Quick start

HTTP streaming

{
"text": "नमस्ते! आपका order 2 दिन में deliver हो जाएगा।",
"language_code": "hi-IN",
"model": "bulbul:v4-flash",
"speaker": "aparna_hi_customer",
"speech_sample_rate": 24000,
"output_audio_codec": "wav",
"pace": 1.0,
"enable_preprocessing": true
}

POST https://api.sarvam.ai/text-to-speech/stream with header api-subscription-key.

WebSocket

Connect to wss://api.sarvam.ai/text-to-speech/ws?model=bulbul:v4-flash with api-subscription-key on the upgrade request. Per utterance, send in order:

{"type": "config", "data": {
"speaker": "aparna_hi_customer",
"language_code": "hi-IN",
"speech_sample_rate": 24000,
"output_audio_codec": "wav",
"pace": 1.0
}}
{"type": "text", "data": {"text": "नमस्ते! मैं आपकी कैसे मदद कर सकती हूँ?"}}
{"type": "flush"}

Audio arrives as {"type":"audio"} messages with base64 in data.audio. Hold the socket open across many calls; send {"type":"ping"} on idle links.


Parameter reference (v4 Flash)

FieldValuesNotes
textstringRequired.
language_codehi-IN, en-IN, bn-IN, gu-IN, kn-IN, ml-IN, mr-IN, od-IN, pa-IN, ta-IN, te-INRequired on REST and HTTP streaming. WebSocket config uses the same field name.
modelbulbul:v4-flashRequired for this guide.
speakerPersona ID, e.g. ritu_hi_customerForm voice_language_style. Unknown values return 400 with the live speaker list. If omitted, default is shubh_enhi_ads.
speech_sample_rate8000, 16000, 22050, 24000 (+ 32000, 44100, 48000 on REST only)Streaming capped at 24 kHz. Default 22050 when omitted.
output_audio_codecwav, mp3, flac, aac (wav, mp3 on WebSocket)Set explicitly on streaming — default is not WAV.
output_audio_bitrate32k – 192kMP3 on HTTP streaming only.
pace0.5 – 2.0, default 1.0Speaking rate.
pitch-0.5 – 0.5, default 0.0Lower/higher while preserving identity.
loudness0.1 – 2.5, default 1.0Output gain — prefer over downstream amplification.
enable_preprocessingbooleanNormalises numbers, dates, mixed script before synthesis.

Values outside allowed ranges return 400. Validate in your client to avoid wasted round trips on live calls.

Tuning guidance

SettingWhere it helps
pace 0.8 – 0.9EdTech, wellness, accessibility, instructional content
pace 1.0Conversational agents — start here
pace 1.1 – 1.2Notifications, news, professional IVR
pitch ±0.1 – 0.2Small brand-tone tweaks; audition larger shifts
loudness 1.0Default unless matching a fixed telephony level

Writing text for natural speech

Input formatting affects perceived quality more than parameter tweaks. Most early issues trace back to this section.

Punctuation for pauses

PunctuationEffect
,Short pause
. / ।Sentence end (use । when the sentence ends in Indic script)
!Emphasis + pause
…Hesitation — use sparingly
Line breakPause between paragraphs

Fillers and code-mixing

For conversational Hinglish, write English in Latin script and Indic words in native script, as a bilingual reader would:

  • "आपका order confirm हो गया है"
  • "Aapka order confirm ho gaya hai"

Use _enhi_ personas (simran_enhi_customer, ishita_enhi_customer, shubh_enhi_banking, …) when copy is genuinely mixed English–Hindi in one utterance, not just a few loan words.

Write each language in its own script

Romanised Hindi/Tamil/etc. degrades quality — the model was trained on native scripts.

LanguageCorrectAvoid
Hindiआपका order confirm हो गया हैAapka order confirm ho gaya hai
Tamilஉங்கள் order deliver ஆகிவிடும்Unga order deliver aayidum

Let preprocessing handle numbers

With enable_preprocessing: true, do not hand-expand numbers, dates, or currency into words — you will usually get a worse reading than the model’s language-specific rules.

Chunk on sentence boundaries

Split long or LLM-generated text at sentence boundaries. Mid-sentence splits strand intonation and sound like an audible seam. Buffer LLM tokens until you have a full sentence, then synthesise.

Match voice to language

Personas are not equally strong in all eleven languages. Use the rankings below and audition cross-language use before production.

Avoid

AvoidWhyFix
Transliterated IndicPoor pronunciationNative script for Indic words
Very long sentencesUnnatural breathingShorter sentences
Overusing …Choppy deliveryCommas or line breaks

Target language code

language_code is required. It drives number, date, and entity reading rules.

LanguageCode
Englishen-IN
Hindihi-IN
Bengalibn-IN
Tamilta-IN
Telugute-IN
Kannadakn-IN
Malayalamml-IN
Marathimr-IN
Gujaratigu-IN
Punjabipa-IN
Odiaod-IN
audio = client.text_to_speech.convert(
text="नमस्ते! Sarvam AI में आपका स्वागत है।",
model="bulbul:v4-flash",
language_code="hi-IN",
speaker="aparna_hi_customer",
)

For mixed copy, set the code to the language whose number and entity conventions you want in speech.


Speaker ID format

Every v4 Flash ID follows voice_language_style:

IDMeaning
ritu_hi_customerRitu · Hindi · customer care
shubh_en_recoveryShubh · English · recovery
simran_enhi_customerSimran · English+Hindi code-mixed · customer care

The same voice can appear under multiple styles (ritu_hi_customer vs ritu_hi_reels) with different energy — switching style is often better than switching voice.


Expert picks for English (en-IN) and Hindi (hi-IN). Treat rank 1 as default; ranks 2–3 are A/B candidates. The full roster is browsable on Voices; the style segment (customer, recovery, sales, …) is a reliable guide when your use case is not listed.

English (en-IN)

Use caseFemaleMale
Sales / lead conversion1. simran_en_sales · 2. aparna_hi_customer1. sunny_en_social · 2. rohan_en_recovery · 3. shubh_hi_ecomm
Collections / payment reminder1. aparna_hi_kyc · 2. sanchita_en_recovery · 3. zarina_en_conversation1. shubh_en_recovery · 2. shubh_hi_ecomm
Customer care & empathetic support1. aparna_en_companion · 2. sanchita_en_companion1. shubh_hi_customer
Appointments / receptionist1. aparna_hi_customer · 2. sanchita_en_insurance—
EdTech / tutor1. aparna_en_edtech—

Hindi (hi-IN)

Use caseFemaleMale
Sales / lead conversion1. aparna_hi_customer1. shubh_hi_customer
Collections / payment reminder1. simran_hi_recovery1. shubh_hi_ecomm · 2. shubh_hi_customer
Customer care & empathetic support1. aparna_hi_customer · 2. aarti_hi_customer—
Appointments / receptionist1. aparna_hi_customer1. shubh_hi_ecomm · 2. ratan_hi_customer_expressive

Roster coverage by language

LanguageVoicesSpeaker IDsNotes
Hindi hi48116Deepest coverage
English en2962Indian accent plus a few other accents
English + Hindi enhi613Code-mixed copy
Marathi mr89
Kannada, Punjabi, Tamil, Telugu3 each4 eachNarrow — audition carefully
Bengali, Gujarati3 / 23 each
Assamese as22Not in the accepted language_code list yet — contact Sarvam before building
Malayalam, Odia——ml-IN and od-IN are valid codes, but the roster has no ml / od personas yet — talk to Sarvam if these are on your roadmap

Decoding API audio (base64)

REST returns base64 in the audios array. Decode before writing to disk.

import base64
from sarvamai import SarvamAI
client = SarvamAI(api_subscription_key="YOUR_SARVAM_API_KEY")
audio = client.text_to_speech.convert(
text="नमस्ते! Sarvam AI में आपका स्वागत है।",
model="bulbul:v4-flash",
language_code="hi-IN",
speaker="aparna_hi_customer",
)
audio_bytes = base64.b64decode("".join(audio.audios))
with open("output.wav", "wb") as f:
f.write(audio_bytes)

For WebSocket streaming, decode each chunk as it arrives. See WebSocket streaming.

Writing raw base64 to a file produces corrupted audio. Always decode first.


Output formats and sample rates

CodecBest forNotes
wavAgents, post-processing, archivalUncompressed
mp3Web, mobile, distributionBitrate configurable on HTTP streaming
aaciOS, adaptive streamingStrong quality at moderate bitrates
flacHigh-fidelity archivalLossless, smaller than WAV
RateAvailable onUse case
8 kHzAllTelephony, IVR, PSTN legs
16 kHzAllVoice agents — common real-time default
22.05 kHzAllGeneral playback
24 kHzAllHighest quality while streaming
32 / 44.1 / 48 kHzREST onlyBroadcast, podcast, studio offline

Defaults to start: agents — wav @ 16 kHz; telephony — wav @ 8 kHz. Do not request 48 kHz if a gateway downsamples to 8 kHz — you only add transfer time.


Handling errors

StatusMeaningAction
200Success—
400Bad speaker, language, or parameter rangeFix request; 400 often lists valid speakers
403Missing/invalid API keyCheck api-subscription-key
429ThrottledBack off with exponential backoff and jitter

On WebSocket, errors often arrive as {"type":"error"} followed by a close — reconnect before retrying.

Log request_id from responses when debugging with support.


WebSocket connection management

  • Idle timeout: connections close after prolonged inactivity — send ping well before that window.
  • Pronunciations: use a pronunciation dictionary for names and brands instead of phonetic spelling hacks.

Key considerations

  • For large numbers, use grouping that matches your locale (e.g. 10,000 or 5,00,000) so preprocessing reads them correctly.
  • v4 Flash does not support SSML — use pace, pitch, loudness, and sentence-boundary chunking for rhythm.
  • Character limits apply per request/session (typically 2,500) — chunk at sentences.

Quick reference

Modelbulbul:v4-flash
StreamingPOST /text-to-speech/stream
RESTPOST /text-to-speech
WebSocketwss://api.sarvam.ai/text-to-speech/ws?model=bulbul:v4-flash
Authapi-subscription-key
Default speaker (if omitted)shubh_enhi_ads
Pace / pitch / loudness1.0 / 0.0 / 1.0
Preprocessingenable_preprocessing: true recommended

Need bulbul:v3 short-name speakers or temperature tuning? See Bulbul v3 best practices.