Best Practice Guide for Bulbul v4 Flash
Bulbul v4 Flash is Sarvam’s low-latency TTS tier for interactive workloads: voice agents, IVR, and anything where a listener is waiting for the first word. This guide applies only to model="bulbul:v4-flash" with persona speaker IDs (voice_language_style). For the stable bulbul:v3 model (short names like shubh, priya), see Bulbul v3 best practices.
Preview voices before you ship: Voices lists every v4 Flash persona with audio samples.
What is Bulbul v4 Flash?
Flash is the interactive tier of the v4 family. It runs on production infrastructure and is available over three transports (HTTP streaming, single-response REST, and WebSocket) so you can match your pipeline instead of forcing one mode.
Choosing the right endpoint
All transports serve the same model; they differ in latency, max sample rate, and how audio is delivered.
If a person is waiting to hear the first word, use HTTP streaming or WebSocket. The single-response endpoint finishes the full clip before replying, which listeners experience as dead air. Reserve REST for offline generation and when you need 32 / 44.1 / 48 kHz.
Reuse connections: Each new HTTPS connection pays DNS, TCP, and TLS before any audio work. Enable keep-alive or pool clients so one connection serves many synthesis calls.
See also Which API to Use.
Quick start
HTTP streaming
POST https://api.sarvam.ai/text-to-speech/stream with header api-subscription-key.
WebSocket
Connect to wss://api.sarvam.ai/text-to-speech/ws?model=bulbul:v4-flash with api-subscription-key on the upgrade request. Per utterance, send in order:
Audio arrives as {"type":"audio"} messages with base64 in data.audio. Hold the socket open across many calls; send {"type":"ping"} on idle links.
Parameter reference (v4 Flash)
Values outside allowed ranges return 400. Validate in your client to avoid wasted round trips on live calls.
Tuning guidance
Writing text for natural speech
Input formatting affects perceived quality more than parameter tweaks. Most early issues trace back to this section.
Punctuation for pauses
Fillers and code-mixing
For conversational Hinglish, write English in Latin script and Indic words in native script, as a bilingual reader would:
-
"आपका order confirm हो गया है" -
"Aapka order confirm ho gaya hai"
Use _enhi_ personas (simran_enhi_customer, ishita_enhi_customer, shubh_enhi_banking, …) when copy is genuinely mixed English–Hindi in one utterance, not just a few loan words.
Write each language in its own script
Romanised Hindi/Tamil/etc. degrades quality — the model was trained on native scripts.
Let preprocessing handle numbers
With enable_preprocessing: true, do not hand-expand numbers, dates, or currency into words — you will usually get a worse reading than the model’s language-specific rules.
Chunk on sentence boundaries
Split long or LLM-generated text at sentence boundaries. Mid-sentence splits strand intonation and sound like an audible seam. Buffer LLM tokens until you have a full sentence, then synthesise.
Match voice to language
Personas are not equally strong in all eleven languages. Use the rankings below and audition cross-language use before production.
Avoid
Target language code
language_code is required. It drives number, date, and entity reading rules.
For mixed copy, set the code to the language whose number and entity conventions you want in speech.
Speaker ID format
Every v4 Flash ID follows voice_language_style:
The same voice can appear under multiple styles (ritu_hi_customer vs ritu_hi_reels) with different energy — switching style is often better than switching voice.
Recommended voices by use case
Expert picks for English (en-IN) and Hindi (hi-IN). Treat rank 1 as default; ranks 2–3 are A/B candidates. The full roster is browsable on Voices; the style segment (customer, recovery, sales, …) is a reliable guide when your use case is not listed.
English (en-IN)
Hindi (hi-IN)
Roster coverage by language
Decoding API audio (base64)
REST returns base64 in the audios array. Decode before writing to disk.
Python
JavaScript
For WebSocket streaming, decode each chunk as it arrives. See WebSocket streaming.
Writing raw base64 to a file produces corrupted audio. Always decode first.
Output formats and sample rates
Defaults to start: agents — wav @ 16 kHz; telephony — wav @ 8 kHz. Do not request 48 kHz if a gateway downsamples to 8 kHz — you only add transfer time.
Handling errors
On WebSocket, errors often arrive as {"type":"error"} followed by a close — reconnect before retrying.
Log request_id from responses when debugging with support.
WebSocket connection management
- Idle timeout: connections close after prolonged inactivity — send
pingwell before that window. - Pronunciations: use a pronunciation dictionary for names and brands instead of phonetic spelling hacks.
Key considerations
- For large numbers, use grouping that matches your locale (e.g.
10,000or5,00,000) so preprocessing reads them correctly. - v4 Flash does not support SSML — use
pace,pitch,loudness, and sentence-boundary chunking for rhythm. - Character limits apply per request/session (typically 2,500) — chunk at sentences.
Quick reference
Need bulbul:v3 short-name speakers or temperature tuning? See Bulbul v3 best practices.