Best Practices for Bulbul v3
This page covers Bulbul v3 only (model="bulbul:v3" or omit model for the API default). For the latest persona-based model, see the Best Practice Guide for Bulbul v4 Flash.
Want to hear v3 speakers? See Voices — Bulbul v3.
Choosing the right API mode
Bulbul v3 supports REST and WebSocket streaming. Picking the right transport has a large impact on latency.
For voice agent pipelines (LLM → TTS), prefer HTTP streaming or WebSocket. REST adds dead air while the full clip is generated.
Pace and temperature
Two parameters control how Bulbul v3 sounds. Start with defaults, then tune for your use case.
Pace (0.5 – 2.0)
pace is speaking rate relative to the voice’s natural speed. Default 1.0.
Temperature (0.01 – 1.0)
temperature controls expressiveness. Default 0.6.
Bulbul v3 does not support pitch or loudness (those apply to bulbul:v2 and bulbul:v4-flash).
Speaker selection by language
Use short-name speakers (shubh, priya, …) with bulbul:v3. Do not use v4 persona IDs with v3.
Varun is a dramatic character voice — not a neutral default. Reserve for thriller or suspense content.
Use-case presets (v3)
Output formats (v3)
For v3-specific limits (SSML, sample rates on streaming), see the Bulbul model page.