Text-to-Speech Rest API
Provides a synchronous REST endpoint where a POST request with text returns base64-encoded audio as response.
The JSON response contains an audios array of base64-encoded WAV strings, not raw binary. Decode before saving or playing:
See TTS best practices for JavaScript and streaming examples.
Common use cases:
- Story narration: Generate expressive audio for audiobooks and narratives
- Podcast generation: Create natural-sounding voiceovers for episodes at scale
- Content creation: Add voice to blogs, articles, and social media posts
- E-learning: Build multilingual course material with clear pronunciation
What You Can Do
Model: Bulbul v3
Bulbul v3 is purpose-built for Indian languages and accents. It handles code-mixed text (e.g., Hinglish), number normalization, and natural prosody out of the box, with minimal preprocessing needed.
Text to Speech Features
Basic Synthesis
Voice Selection
Advanced Options
Basic Text to Speech Synthesis
Convert text to natural-sounding speech with high quality. Features include:
- Multiple voice options
- Support for Indian languages
- Natural prosody and intonation
- High-quality audio output
API Response Format
Supported audio formats: WAV (default), MP3, Linear16, Mulaw, Alaw, Opus, FLAC, AAC
Decoding Audio Examples
Python:
JavaScript:
Error Responses
All errors return a JSON object with an error field (message, code, request_id). The full error-code table, retry guidance, and SDK exception reference live on the central Errors & Troubleshooting page.
Errors specific to this endpoint:
Error Handling Code Example
Check out our detailed API Reference to explore Text to Speech and all available options.
Need help? Contact us on discord for guidance.