Stream Speech (SSE)
Synthesize speech and stream the audio back over Server-Sent Events. Same body as /waves/v1/tts — the only difference is the response is a stream of base64-encoded PCM chunks instead of one binary blob.
Pick the model with the model body parameter, same as the sync route.
The same URL serves the WebSocket endpoint. wss://api.smallest.ai/waves/v1/tts/live accepts a WebSocket upgrade for streaming-text scenarios (LLM token streams, live captioning). The HTTP POST documented on this page returns SSE; use wss:// to use the WebSocket protocol instead. See the WebSocket reference.
When to use this
- Use this when you want playback to start before synthesis is complete — long passages, latency-sensitive UI, live narration.
- Use sync
/waves/v1/ttswhen total latency doesn’t matter and you’d rather get one buffer. - Use
/waves/v1/tts/live(WebSocket) when the text arrives incrementally (LLM token stream). SSE assumes you have the full text up front.
How it works
- POST your text + voice settings — same payload as
/waves/v1/tts, plus optionalmodel. - The response is
Content-Type: text/event-stream. Each chunk frame isevent: audio\nfollowed bydata: {"audio": "<base64-pcm>", "done": false, "status": "206"}\n\n. - Decode each chunk’s
audiofield with base64 and feed the PCM bytes to your audio pipeline (browserMediaSource, ffmpeg pipe, raw PCM player, etc.). - A final
data: {"status": "200", "done": true}\n\nframe marks end of stream. Detect the terminator withdone == true; every chunk frame also carriesdone: false, so"done" in msgmatches every frame.
Examples
cURL
Common gotchas
- Use a streaming-friendly client.
curl -N, Pythoniter_lines, or afetchReadableStreamreader. Buffering clients will hide the latency win. - Audio is base64 inside the event payload, not the raw event bytes. Decode the
data.audiofield per event. output_format=pcmgives the lowest overhead for streaming playback.wav/mp3work but add per-chunk framing bytes.
Authentication
API key authentication. Include your key as Authorization: Bearer YOUR_API_KEY. Access tokens are not accepted on this endpoint.
Headers
Enterprise plans only. Opt in if you want this request's content deleted after 7 days. Omit it to retain content, which is the default.
Request
The text to convert to speech. Max 8000 characters after trim; whitespace-only strings are rejected.
TTS model to route the request to. Controls which model pool serves this synthesis.
lightning_v3.1(default) — standard Lightning v3.1.lightning_v3.1_pro— Lightning v3.1 Pro pool. Improved audio quality and naturalness, with a curated voice catalog. See the Lightning v3.1 Pro model card for supported voice IDs.
Same concurrency and latency profile across both. Other request parameters behave identically.
Language code for synthesis. Influences pronunciation, number/date normalization, and phoneme selection.
Default on lightning_v3.1_pro: when language is omitted, the
Pro pool defaults to en + hi (mixed Indian + Western English
coverage, auto-detected from the input text).
Each voice has its own tags.language set in the voice catalog —
query GET /waves/v1/lightning-v3.1/get_voices. Pass a language
the voice was trained on; passing other codes is accepted by the
API but produces English-pronounced output.
auto (recommended for cross-language use cases): routes internally
based on the input text. Any English or Hindi voice can be used
across all supported languages when auto is set; the platform
handles language-appropriate routing without needing a code per
call.
On lightning_v3.1 — 20 supported languages:
- 10 European: English, Spanish, French, German, Italian, Dutch, Swedish, Portuguese, Polish, Russian
- 10 Indic: Hindi, Marathi, Gujarati, Punjabi, Bengali, Odia, Tamil, Telugu, Kannada, Malayalam
On lightning_v3.1_pro — 31 supported languages (adds 11 over base):
- 13 European: base 10 plus Greek, Finnish, Norwegian
- 8 Asian & Middle Eastern: Chinese, Japanese, Korean, Indonesian, Malay, Vietnamese, Turkish, Arabic
- 10 Indic: same as base
- Pass
en→ UK + American accented English. - Pass
hi→ Indian accented English + Hindi (code-switching). - Omit
language→ defaults toen + hi(mixed Indian + Western English coverage, auto-detected from input text).
Optional. Sets the language used to read out numeric content — numbers, currency amounts, times, and the numeric parts of dates and years — independently of the synthesis voice. Ordinary words are not translated.
- If you omit
language, this value also becomes the synthesis language: model selection and voice routing follow it. - If you set
languageexplicitly,languagealways wins for synthesis andnumber_pronunciation_languageonly changes how numeric content is normalized. It works both ways — read numbers in Hindi under an English voice, or in English under a Hindi voice (tuned for Indian, often mixed-script, use cases). - Omit this field to keep the existing behaviour — normalization
follows
language.
Note: only numeric tokens are re-spoken; the words around them stay in the text language. On a cross-language request names may also render in the target script (e.g. "Smith" → "स्मिथ"), which is generally the desired reading for native-language voices.
Accepts the same language codes as language (including auto,
nl, sv).
Opt-in flag that reads digit-flanked math operators (5 x 3,
2 ^ 10, 6 ÷ 2) as words instead of leaving them for the
default number reader. Off by default because in real traffic
digit-flanked NxN is more often a product dimension, the
24x7 idiom, or a vehicle-registration code than an actual
multiplication.
When true, the normalizer replaces the operator with the
spoken word matched to number_pronunciation_language:
Localized only for hi and mr; every other language falls
back to the English words. The operator word follows
number_pronunciation_language, not the synthesis
language, so language=en, number_pronunciation_language=hi
reads 6 x 7 as "छः गुणा सात".
Matching rules: unambiguous glyphs (× ÷ * ^ ** = + and the
wrong-glyph x/X) fire glued or spaced (5x3, 5 x 3).
The ambiguous - – − and / fire only when
space-padded, so 5-3 stays a range and 1/2 stays a
fraction. See Math notation
for the full lexicon, known limitations (product dimensions,
24x7 idiom, vehicle-reg codes), and EU-language
localizations.
Opt-in profanity filter for the submitted text. Off by default; passing this object is the only way to turn it on. The filter never rewrites your text — it either lets the request through or rejects it before synthesis.
action: "reject" returns HTTP 400 with error_code: "CONTENT_FILTER_BLOCKED", the language checked and a
match_count; the matched terms are never returned or logged.
action: "flag" synthesizes normally and records the match.
Matching is whole-word, not substring. If no verdict is returned the request fails open and the audio is synthesized unfiltered.
See Content filter.
Format of the returned audio. pcm is the lowest-latency option
but requires a decoder to play; mp3 and wav are directly
playable in browsers and most media players. The server default
is pcm when the field is omitted — the API playground uses
mp3 so the generated audio is directly playable.
The IDs of the pronunciation dictionaries to use for speech generation. Available on both lightning_v3.1 and lightning_v3.1_pro.
WebSocket-only feature. Accepted on this endpoint but ignored — no per-word timing information is returned in the sync HTTP or SSE response shape. To receive status: "word_timestamp" frames with per-word { id, word, start, end } data, use the WebSocket endpoint wss://api.smallest.ai/waves/v1/tts/live. See Word-level timestamps.
Optional client-provided session identifier for correlation. Only alphanumeric characters, hyphens, underscores, and dots are allowed. Max 128 characters. Echoed back in response headers as X-External-Session-Id.
Optional client-provided request identifier for correlation. Only alphanumeric characters, hyphens, underscores, and dots are allowed. Max 128 characters. Echoed back in response headers as X-External-Request-Id.
Response headers
Internal session identifier (system-generated UUID).
Internal request identifier (system-generated UUID).
Echoed client-provided session_id (empty if not provided).
Echoed client-provided request_id (empty if not provided).