Skip to navigation

Stream Speech (WebSocket)

View as Markdown

Real-time text-to-speech over a persistent WebSocket connection. The model field in the request payload selects which Lightning pool serves the synthesis.

When to use this

  • Use this when text arrives incrementally (LLM token streams, live captioning, conversational pipelines where playback should start as soon as the first chunk is ready).
  • POST to /waves/v1/tts/live (SSE) when you have the full text up front but still want chunked playback. (Same URL, different protocol — HTTP POST gets you SSE; WSS connect gets you WebSocket.)
  • Use /waves/v1/tts (sync) when total latency doesn’t matter.

Selecting the model

Pass "model": "lightning_v3.1" (default) or "model": "lightning_v3.1_pro" on each request. Concurrency and latency are identical across both. Voice catalogs differ — see the Lightning v3.1 and Lightning v3.1 Pro model cards for the per-model catalog.

Language behaviour

auto (recommended for cross-language use cases): routes internally based on the input text. Any English or Hindi voice can be used across all supported languages when auto is set.

On lightning_v3.1 — 20 accepted language codes (10 European + 10 Indic). The trained voice catalog covers 12 of these directly; the other 8 route through English or Hindi voices.

On lightning_v3.1_pro — 31 languages with dedicated voices (10 Indic, 8 Asian & Middle Eastern, 13 European including Dutch and Swedish):

  • Pass language: en → UK + American accented English.
  • Pass language: hi → Indian accented English + Hindi (code-switching).
  • Pass the ISO 639-1 code of any other Pro language (e.g. ta, de, ja) with a matching Pro voice. See the Lightning v3.1 Pro model card for the full list.
  • Omit language → defaults to en + hi (mixed Indian + Western English coverage).

Optional features

Set word_timestamps: true to receive per-word timing events interleaved with the audio chunks (status: "word_timestamp"). Supported on English + Hindi base-queue voices. See Word-level timestamps.

Send the same context_id on a sequence of text fragments to have them buffered, joined at natural sentence boundaries, and spoken as one continuous generation instead of one reset-per-fragment. See Continuations.

Connection timeout

The server closes idle WebSocket connections to free resources. The default idle timeout is 60 seconds — if your client does not send a message within that window the server closes the connection with:

{"status": "error", "message": "Connection timed out after 60 seconds of inactivity"}

Override the value with the timeout query parameter on the URL:

wss://api.smallest.ai/waves/v1/tts/live?timeout=120

Pass a positive integer (seconds). Smaller values are honored verbatim (e.g. ?timeout=5 closes after 5 s of silence); larger values are clamped to the maximum of 180 seconds. Use a larger value when your application has known pauses between turns — voice agents with long human-thinking windows, agentic pipelines waiting on an LLM round-trip, etc.

The timeout is reset on every message you send (binary audio in, JSON control in).

Keep-alive

To hold a connection open past the timeout without sending audio, send a JSON keep-alive frame:

{"type": "ping"}

The server replies {"type": "pong"} and resets the inactivity timer. keepalive, keep_alive, and keep-alive are accepted as aliases. The frame is handled at the API edge — it never reaches the model and does not affect synthesis. Prefer this over a raised timeout: the pong also tells you the connection is still live, and it is the only way to stay open past the 180-second cap.

Migrating from /waves/v1/lightning-v3.1/get_speech/stream

Same protocol, same payload shape — only the URL changes. Existing clients should:

  1. Update the WebSocket URL to wss://api.smallest.ai/waves/v1/tts/live.
  2. Optionally add "model": "lightning_v3.1_pro" to route to the Pro pool. Omitting model keeps the existing standard-pool behavior.

Voice IDs, sample rates, auth, and the response/streaming format are unchanged, so downstream audio handling, jitter buffers, and barge-in logic stay the same.

Handshake

WSS
wss://api.smallest.ai/waves/v1/tts/live

Authentication

AuthorizationBearer

API key authentication. Include your key as Authorization: Bearer YOUR_API_KEY. Access tokens are not accepted on this endpoint.

Send

TtsRequestobjectRequired
Send a JSON message with `voice_id`, `text`, and optional parameters (including `model`) to generate speech audio.

Receive

TtsResponseobjectRequired
Receive audio data chunks and completion status from the server.