HTTP Streaming API
POST /voices/clone/stream. Send text and a saved voice_id, get a binary audio stream back. The response starts arriving as soon as the first sentence is ready, so you can begin playback without waiting for the full file.
No WebSocket handshake, no config messages, no connection lifecycle. One HTTP request, one streamed response.
Common use cases:
- Backend audio generation: Pipe audio directly to a file, S3, or a downstream service
- API proxying: Forward the stream to your frontend or mobile client as-is
- Serverless / edge: Works in any environment that supports HTTP, no WebSocket runtime needed
HTTP Stream vs WebSocket: When to Use Which
Both give you streaming audio from a cloned voice. The difference is how much control you need.
Use HTTP Stream when:
- You have a complete text and just need audio back
- You’re generating audio server-side (batch jobs, API endpoints, CI pipelines)
- Your runtime doesn’t support WebSocket
- You want the simplest possible integration —
curl -Nworks out of the box
Use WebSocket when:
- You’re building a conversational agent that streams text incrementally (e.g., from an LLM)
- You need to send multiple texts without reconnecting
- Low time-to-first-byte on successive utterances matters (connection is already warm)
See the WebSocket streaming guide for the socket protocol.
Code Examples
Create a voice with POST /voices/create first, then pass the returned voice_id.
Runnable samples are also on the REST Stream endpoint in the API Reference.
Request fields
Same multipart contract as POST /voices/clone, with these differences:
Codec guidance
How a client should handle the stream
- Do not parse JSON. A successful response body is raw audio bytes.
- Check status before reading the body. Errors before the first audio byte return normal HTTP error JSON (
400,402,403,422,429,503). - After the first byte, the stream is binary. Mid-stream failures end the connection; there is no in-band error envelope.
- Progressive playback: pipe chunks into a decoder or a media player that accepts incremental MP3/PCM.
- Disconnect: closing the client connection cancels the stream; generation work already done may still be billed.