Stream Speech (WebSocket)
Real-time text-to-speech over a persistent WebSocket connection. The
model field in the request payload selects which Lightning pool serves
the synthesis.
When to use this
- Use this when text arrives incrementally (LLM token streams, live captioning, conversational pipelines where playback should start as soon as the first chunk is ready).
- POST to
/waves/v1/tts/live(SSE) when you have the full text up front but still want chunked playback. (Same URL, different protocol — HTTP POST gets you SSE; WSS connect gets you WebSocket.) - Use
/waves/v1/tts(sync) when total latency doesn’t matter.
Selecting the model
Pass "model": "lightning_v3.1" (default) or
"model": "lightning_v3.1_pro" on each request. Concurrency and latency
are identical across both. Voice catalogs differ — see the
Lightning v3.1 and
Lightning v3.1 Pro
model cards for the per-model catalog.
Language behaviour
auto (recommended for cross-language use cases): routes internally
based on the input text. Any English or Hindi voice can be used across
all supported languages when auto is set.
On lightning_v3.1 — 20 accepted language codes (10 European + 10
Indic). The trained voice catalog covers 12 of these directly; the
other 8 route through English or Hindi voices.
On lightning_v3.1_pro — 31 languages with dedicated voices (10
Indic, 8 Asian & Middle Eastern, 13 European including Dutch and
Swedish):
- Pass
language: en→ UK + American accented English. - Pass
language: hi→ Indian accented English + Hindi (code-switching). - Pass the ISO 639-1 code of any other Pro language (e.g.
ta,de,ja) with a matching Pro voice. See the Lightning v3.1 Pro model card for the full list. - Omit
language→ defaults toen + hi(mixed Indian + Western English coverage).
Optional features
Set word_timestamps: true to receive per-word timing events
interleaved with the audio chunks (status: "word_timestamp").
Supported on English + Hindi base-queue voices. See
Word-level timestamps.
Send the same context_id on a sequence of text fragments to have
them buffered, joined at natural sentence boundaries, and spoken as
one continuous generation instead of one reset-per-fragment. See
Continuations.
Connection timeout
The server closes idle WebSocket connections to free resources. The default idle timeout is 60 seconds — if your client does not send a message within that window the server closes the connection with:
Override the value with the timeout query parameter on the URL:
Pass a positive integer (seconds). Smaller values are honored
verbatim (e.g. ?timeout=5 closes after 5 s of silence); larger
values are clamped to the maximum of 180 seconds. Use a larger
value when your application has known pauses between turns — voice
agents with long human-thinking windows, agentic pipelines waiting on
an LLM round-trip, etc.
The timeout is reset on every message you send (binary audio in, JSON control in).
Keep-alive
To hold a connection open past the timeout without sending audio, send a JSON keep-alive frame:
The server replies {"type": "pong"} and resets the inactivity timer.
keepalive, keep_alive, and keep-alive are accepted as aliases.
The frame is handled at the API edge — it never reaches the model and
does not affect synthesis. Prefer this over a raised timeout: the
pong also tells you the connection is still live, and it is the only
way to stay open past the 180-second cap.
Migrating from /waves/v1/lightning-v3.1/get_speech/stream
Same protocol, same payload shape — only the URL changes. Existing clients should:
- Update the WebSocket URL to
wss://api.smallest.ai/waves/v1/tts/live. - Optionally add
"model": "lightning_v3.1_pro"to route to the Pro pool. Omittingmodelkeeps the existing standard-pool behavior.
Voice IDs, sample rates, auth, and the response/streaming format are unchanged, so downstream audio handling, jitter buffers, and barge-in logic stay the same.
Handshake
Authentication
API key authentication. Include your key as Authorization: Bearer YOUR_API_KEY. Access tokens are not accepted on this endpoint.