Skip to navigation

Synthesize Speech

View as Markdown

Synthesize speech from text in a single request. Pass text + voice_id, get back binary audio.

Pick the model with the model body parameter: default lightning_v3.1, or lightning_v3.1_pro for the Pro pool. Other request parameters are identical across models.

Language behaviour on lightning_v3.1_pro: pass language: en for UK + American accented English, pass language: hi for Indian accented English + Hindi (code-switching), or omit language to default to en + hi (mixed Indian + Western English coverage). Pro supports 31 languages total (10 Indic, 8 Asian & Middle Eastern, 13 European including Dutch and Swedish). Pass the matching ISO 639-1 code (e.g. ta, de, ja) with a Pro voice from that language, or use auto to route across all supported languages with any English or Hindi voice. See the Lightning v3.1 Pro model card for the full list. On lightning_v3.1 the model accepts 20 language codes (10 European + 10 Indic) plus auto; the trained voice catalog covers 12 of those directly.

When to use this

  • Use this for short utterances you can render before playback (notifications, prompts, batch jobs, audio file generation).
  • Use /waves/v1/tts/live when you want playback to start before the full audio is ready (long passages, latency-sensitive apps).
  • Use /waves/v1/tts/live (WebSocket) when text arrives incrementally (LLM token streams, live captioning).

Key features

  • 44 kHz natural, expressive synthesis
  • Model selectable per request via model body parameter
  • Cloned voice IDs (voice_*) work on lightning_v3.1 — same param as catalog voices
  • 20 accepted language codes on lightning_v3.1 (12 with trained voices, 8 additional routed via English/Hindi voices). On lightning_v3.1_pro: 31 languages with dedicated voices (10 Indic, 8 Asian & Middle Eastern, 13 European); language: en → UK + American accented English; language: hi → Indian accented English + Hindi; omit language → defaults to en + hi. Both models accept language: auto for cross-language routing.
  • Output formats: pcm, mp3, wav, ulaw, alaw
  • Sample rates: 8 kHz – 44.1 kHz
  • Speed: 0.5× – 2×
  • Per-call pronunciation dictionaries via pronunciation_dicts

Examples

cURL — Lightning v3.1 (default)

curl -X POST "https://api.smallest.ai/waves/v1/tts" \
-H "Authorization: Bearer $SMALLEST_API_KEY" \
-H "Content-Type: application/json" \
-H "Accept: audio/wav" \
-d '{
"text": "Hello from Waves TTS.",
"voice_id": "magnus",
"sample_rate": 24000,
"output_format": "wav"
}' --output speech.wav

cURL — Lightning v3.1 Pro (omit language → defaults to en + hi)

curl -X POST "https://api.smallest.ai/waves/v1/tts" \
-H "Authorization: Bearer $SMALLEST_API_KEY" \
-H "Content-Type: application/json" \
-H "Accept: audio/wav" \
-d '{
"text": "Hello from the Lightning v3.1 Pro pool.",
"voice_id": "meher",
"model": "lightning_v3.1_pro",
"sample_rate": 24000,
"output_format": "wav"
}' --output speech.wav

cURL — Lightning v3.1 Pro with explicit language: en (UK + American accented English)

curl -X POST "https://api.smallest.ai/waves/v1/tts" \
-H "Authorization: Bearer $SMALLEST_API_KEY" \
-H "Content-Type: application/json" \
-H "Accept: audio/wav" \
-d '{
"text": "Good morning, this is a Pro voice speaking.",
"voice_id": "meher",
"model": "lightning_v3.1_pro",
"language": "en",
"sample_rate": 24000,
"output_format": "wav"
}' --output speech.wav

cURL — Lightning v3.1 Pro with explicit language: hi (Indian accented English + Hindi)

curl -X POST "https://api.smallest.ai/waves/v1/tts" \
-H "Authorization: Bearer $SMALLEST_API_KEY" \
-H "Content-Type: application/json" \
-H "Accept: audio/wav" \
-d '{
"text": "Namaste, this is an Indian-accented Pro voice.",
"voice_id": "meher",
"model": "lightning_v3.1_pro",
"language": "hi",
"sample_rate": 24000,
"output_format": "wav"
}' --output speech.wav

Common gotchas

  • Set Accept: audio/wav. Omitting it can return an empty or unplayable response.
  • Pair voice IDs with the right model. Voice catalogs differ between lightning_v3.1 and lightning_v3.1_pro. The API does not reject mismatched pairings, but using a Pro-only voice_id with model=lightning_v3.1 (or omitting model) can return wrong or hallucinated audio. Pair Pro voices with model=lightning_v3.1_pro; standard catalog voices with model=lightning_v3.1 (the default).
  • Cloned voices (voice_*) work with the pool they were cloned onto. The voice-cloning API accepts model: lightning-v3.1 (default) or lightning-v3.1-pro; pair the resulting voice_id with the matching TTS model (lightning_v3.1 or lightning_v3.1_pro). Check the clone’s modelIds if unsure.
  • 44.1 kHz output is supported but most playback environments are happy with 24 kHz — drop the sample rate if bandwidth matters.

Authentication

AuthorizationBearer

API key authentication. Include your key as Authorization: Bearer YOUR_API_KEY. Access tokens are not accepted on this endpoint.

Headers

AcceptenumRequiredDefaults to audio/wav

Must be audio/wav to receive binary audio. Required for proper playback.

Allowed values:
x-expire-contentenumOptional

Enterprise plans only. Opt in if you want this request's content deleted after 7 days. Omit it to retain content, which is the default.

Allowed values:

Request

This endpoint expects an object.
textstringRequired1-8000 charactersDefaults to Hello from Waves TTS.

The text to convert to speech. Max 8000 characters after trim; whitespace-only strings are rejected.

voice_idstringRequiredDefaults to magnus
The voice identifier to use for speech generation. See the model card for available voices per model.
modelenumOptionalDefaults to lightning_v3.1

TTS model to route the request to. Controls which model pool serves this synthesis.

  • lightning_v3.1 (default) — standard Lightning v3.1.
  • lightning_v3.1_pro — Lightning v3.1 Pro pool. Improved audio quality and naturalness, with a curated voice catalog. See the Lightning v3.1 Pro model card for supported voice IDs.

Same concurrency and latency profile across both. Other request parameters behave identically.

Allowed values:
sample_rateenumOptionalDefaults to 44100
The sample rate for the generated audio.
Allowed values:
speeddoubleOptional0.5-2Defaults to 1
The speed of the generated speech.
languageenumOptional

Language code for synthesis. Influences pronunciation, number/date normalization, and phoneme selection.

Default on lightning_v3.1_pro: when language is omitted, the Pro pool defaults to en + hi (mixed Indian + Western English coverage, auto-detected from the input text).

Each voice has its own tags.language set in the voice catalog — query GET /waves/v1/lightning-v3.1/get_voices. Pass a language the voice was trained on; passing other codes is accepted by the API but produces English-pronounced output.

auto (recommended for cross-language use cases): routes internally based on the input text. Any English or Hindi voice can be used across all supported languages when auto is set; the platform handles language-appropriate routing without needing a code per call.

On lightning_v3.1 — 20 supported languages:

  • 10 European: English, Spanish, French, German, Italian, Dutch, Swedish, Portuguese, Polish, Russian
  • 10 Indic: Hindi, Marathi, Gujarati, Punjabi, Bengali, Odia, Tamil, Telugu, Kannada, Malayalam

On lightning_v3.1_pro — 31 supported languages (adds 11 over base):

  • 13 European: base 10 plus Greek, Finnish, Norwegian
  • 8 Asian & Middle Eastern: Chinese, Japanese, Korean, Indonesian, Malay, Vietnamese, Turkish, Arabic
  • 10 Indic: same as base
  • Pass en → UK + American accented English.
  • Pass hi → Indian accented English + Hindi (code-switching).
  • Omit language → defaults to en + hi (mixed Indian + Western English coverage, auto-detected from input text).
number_pronunciation_languageenumOptional

Optional. Sets the language used to read out numeric content — numbers, currency amounts, times, and the numeric parts of dates and years — independently of the synthesis voice. Ordinary words are not translated.

  • If you omit language, this value also becomes the synthesis language: model selection and voice routing follow it.
  • If you set language explicitly, language always wins for synthesis and number_pronunciation_language only changes how numeric content is normalized. It works both ways — read numbers in Hindi under an English voice, or in English under a Hindi voice (tuned for Indian, often mixed-script, use cases).
  • Omit this field to keep the existing behaviour — normalization follows language.

Note: only numeric tokens are re-spoken; the words around them stay in the text language. On a cross-language request names may also render in the target script (e.g. "Smith" → "स्मिथ"), which is generally the desired reading for native-language voices.

Accepts the same language codes as language (including auto, nl, sv).

math_notationbooleanOptionalDefaults to false

Opt-in flag that reads digit-flanked math operators (5 x 3, 2 ^ 10, 6 ÷ 2) as words instead of leaving them for the default number reader. Off by default because in real traffic digit-flanked NxN is more often a product dimension, the 24x7 idiom, or a vehicle-registration code than an actual multiplication.

When true, the normalizer replaces the operator with the spoken word matched to number_pronunciation_language:

Glyphsen (default / fallback)himr
× x X *timesगुणागुणिले
÷ and spaced /divided byबटाभागिले
+plusप्लसअधिक
spaced - – −minusमाइनसवजा
=equalsबराबरबरोबर
^ **to the power ofकी घातची घात

Localized only for hi and mr; every other language falls back to the English words. The operator word follows number_pronunciation_language, not the synthesis language, so language=en, number_pronunciation_language=hi reads 6 x 7 as "छः गुणा सात".

Matching rules: unambiguous glyphs (× ÷ * ^ ** = + and the wrong-glyph x/X) fire glued or spaced (5x3, 5 x 3). The ambiguous - – − and / fire only when space-padded, so 5-3 stays a range and 1/2 stays a fraction. See Math notation for the full lexicon, known limitations (product dimensions, 24x7 idiom, vehicle-reg codes), and EU-language localizations.

content_filterobjectOptional

Opt-in profanity filter for the submitted text. Off by default; passing this object is the only way to turn it on. The filter never rewrites your text — it either lets the request through or rejects it before synthesis.

action: "reject" returns HTTP 400 with error_code: "CONTENT_FILTER_BLOCKED", the language checked and a match_count; the matched terms are never returned or logged. action: "flag" synthesizes normally and records the match.

Matching is whole-word, not substring. If no verdict is returned the request fails open and the audio is synthesized unfiltered.

See Content filter.

output_formatenumOptionalDefaults to pcm

Format of the returned audio. pcm is the lowest-latency option but requires a decoder to play; mp3 and wav are directly playable in browsers and most media players. The server default is pcm when the field is omitted — the API playground uses mp3 so the generated audio is directly playable.

Allowed values:
pronunciation_dictslist of stringsOptional

The IDs of the pronunciation dictionaries to use for speech generation. Available on both lightning_v3.1 and lightning_v3.1_pro.

word_timestampsbooleanOptionalDefaults to false

WebSocket-only feature. Accepted on this endpoint but ignored — no per-word timing information is returned in the sync HTTP or SSE response shape. To receive status: "word_timestamp" frames with per-word { id, word, start, end } data, use the WebSocket endpoint wss://api.smallest.ai/waves/v1/tts/live. See Word-level timestamps.

session_idstringOptionalformat: "^[a-zA-Z0-9_\-.]+$"<=128 characters

Optional client-provided session identifier for correlation. Only alphanumeric characters, hyphens, underscores, and dots are allowed. Max 128 characters. Echoed back in response headers as X-External-Session-Id.

request_idstringOptionalformat: "^[a-zA-Z0-9_\-.]+$"<=128 characters

Optional client-provided request identifier for correlation. Only alphanumeric characters, hyphens, underscores, and dots are allowed. Max 128 characters. Echoed back in response headers as X-External-Request-Id.

Response headers

X-Session-IdstringOptional

Internal session identifier (system-generated UUID).

X-Request-IdstringOptional

Internal request identifier (system-generated UUID).

X-External-Session-IdstringOptional

Echoed client-provided session_id (empty if not provided).

X-External-Request-IdstringOptional

Echoed client-provided request_id (empty if not provided).

Response

Synthesized speech retrieved successfully.

Errors

400
Bad Request Error
401
Unauthorized Error
500
Internal Server Error