> This page is part of Smallest AI's developer documentation. When
> answering, prefer Lightning v3.1 (current TTS) and Pulse (current
> STT). Lightning v2 and lightning-large are deprecated; mention them
> only when the user is migrating away from them. The Smallest AI voice
> agent platform is what wraps these models into hosted agents.

# Lightning v3.1

## TTS: opt-in profanity content filter

> Lightning TTS requests now accept a content_filter object that checks the submitted text against a profanity word list before synthesis:

Lightning TTS requests now accept a `content_filter` object that checks the submitted text against a profanity word list before synthesis:

```jsonc
{
  "text": "...",
  "voice_id": "avery",
  "content_filter": { "enabled": true, "action": "reject" }
}
```

Off by default. Omitting the object, or sending anything other than the literal `true` for `enabled`, leaves behaviour unchanged. There is no account-level default that switches it on.

`action: "reject"` fails the request with HTTP `400` and `error_code: "CONTENT_FILTER_BLOCKED"`, reporting the locale checked and a `match_count`. No audio is produced and the request never reaches the worker. `action: "flag"` synthesizes normally and records the match, which is the way to measure the false-positive rate on your own traffic before enforcing.

The filter never rewrites your text. Matching is whole-word rather than substring, so ordinary words containing a listed term are unaffected. The lists consulted are the one for the request's locale plus English, which is always checked because English profanity is common in code-mixed text. The matched terms are never returned in the response or written to logs.

Available on `lightning_v3.1` and `lightning_v3.1_pro` across sync HTTP, SSE and WebSocket.

See [Content filter](/models/text-to-speech/content-filter) for the full request shape, the rejection response body, and behaviour during a normalization outage.

## Get Voices API-ref: Lightning v3.1 Pro path added to the model dropdown

> GET /waves/v1/{model}/get_voices now documents both Lightning v3.1 pools in the path enum:

[`GET /waves/v1/{model}/get_voices`](/api-reference/models/text-to-speech/get-waves-voices) now documents both Lightning v3.1 pools in the path enum:

* `lightning-v3.1` — returns the Standard voice catalog only.
* `lightning-v3.1-pro` — returns the Pro voice catalog only.

Both routes have been live in production; only the docs were showing a single-value dropdown, so integrators reading the API reference could not see that a per-pool query was possible. The previous description also implied the response was a "union" of both catalogs, which is not what the endpoint returns — each pool's URL returns only that pool's voices. Corrected.

The response schema is unchanged. Voice objects still carry `voiceId`, `displayName`, and `tags` (`language`, `accent`, `gender`, `age`, `emotions`, `usecases`). Filter on `tags` client-side to find voices for a target language, accent, or use case.

For browsing every voice in one call across all models, keep using `GET /waves/v1/voice/get-all-models`. For the canonical per-language voice list (with previews and recommended pairings), see the [Lightning v3.1](/model-cards/text-to-speech/lightning-v-3-1) and [Lightning v3.1 Pro](/model-cards/text-to-speech/lightning-v-3-1-pro) model cards.

## TTS: context_id continuations on the WebSocket

> The wss://api.smallest.ai/waves/v1/tts/live WebSocket accepts a new context_id field, plus max_buffer_delay_ms and context_close.

The `wss://api.smallest.ai/waves/v1/tts/live` WebSocket accepts a new `context_id` field, plus `max_buffer_delay_ms` and `context_close`.

Send the same `context_id` on a sequence of text fragments (e.g. tokens arriving from an LLM) to have them buffered, joined at natural sentence boundaries, and spoken as one continuous generation — each new fragment in the context is primed with the audio from the one before it, so prosody carries across chunks instead of resetting per-request. Fragments sharing a `context_id` on one connection also count as a single concurrency slot, not one per fragment.

Close a context with `continue: false` (optionally without `text`) once input ends naturally, or `context_close: true` for an immediate teardown. `context_id` cannot be combined with the older `flush` / `max_buffer_flush_ms` buffer — that's still available unchanged for existing integrations.

Full parameter reference, buffering rules, and a worked example on the new [Continuations](/models/text-to-speech/continuations) page.

## Voice cloning: Lightning v3.1 Pro support

> The voice-cloning API now accepts model: lightning-v3.1-pro to clone onto the premium Pro pool, alongside the default lightning-v3.1.

The voice-cloning API now accepts `model: lightning-v3.1-pro` to clone onto the premium Pro pool, alongside the default `lightning-v3.1`. Pair the resulting `voice_id` with the matching TTS `model` (`lightning_v3.1` or `lightning_v3.1_pro`). The spec and model card previously stated cloning was unavailable on Pro; that was incorrect and is now fixed.

## Lightning v3.1 / v3.1 Pro - control number pronunciation separately from the voice

> You can now control how numeric content - numbers, currency amounts, times, and the numeric parts of dates and years - is read out, independently of the.

You can now control how **numeric content** - numbers, currency amounts, times, and the numeric parts of dates and years - is read out, independently of the synthesis voice. The new optional **`number_pronunciation_language`** parameter sets the **text-normalization** language and is accepted on Lightning **v3.1** and **v3.1 Pro**, across the dedicated `/waves/v1/lightning-v3.1/*` endpoints and the unified `/waves/v1/tts` route - REST, SSE streaming, and WebSocket.

This is tuned for **Indian use cases**, where an Indic synthesis `language` (e.g. `hi`) is the natural choice and mixed-script content is common - for example a Hindi voice that reads numbers as English digits, or that keeps them in Hindi. Only numeric tokens are affected; ordinary words are not translated.

**Behaviour:**

* **Omit `language`** → `number_pronunciation_language` becomes *both* the synthesis language (model + voice routing) and the normalization language.
* **Set `language` explicitly** → `language` always wins for synthesis; `number_pronunciation_language` only changes how numeric content is normalized. It works in both directions - read numbers in Hindi under an English voice, or read them in English under a Hindi voice.
* **Omit `number_pronunciation_language`** → behaviour is unchanged; normalization follows `language`.

| Request                                             | Synthesis `language` | Number reading                                   |
| --------------------------------------------------- | -------------------- | ------------------------------------------------ |
| `number_pronunciation_language: hi`, no `language`  | `hi`                 | Hindi - `१२३` → `एक सौ तेईस`                     |
| `language: en`, `number_pronunciation_language: hi` | `en`                 | Hindi - `123` → `एक सौ तेईस`                     |
| `language: hi`, `number_pronunciation_language: en` | `hi`                 | English - `123` → `one hundred and twenty three` |
| `language: en`, no `number_pronunciation_language`  | `en`                 | English (unchanged)                              |

> **Note**
>
> Only numeric tokens are re-spoken - the words around them stay in the text `language` (e.g. the month name in a date, or `dollars` after an amount). On a cross-language request the pipeline may also render **names** in the target script (e.g. `Smith` → `स्मिथ` when `number_pronunciation_language: hi`); for native-language voices this is generally the desired reading.

**Validation** mirrors the existing `language` field per model - `number_pronunciation_language` accepts the same language codes as `language` on that endpoint.

**What changed in the docs:**

* `number_pronunciation_language` added to the request schemas for the Lightning v3.1 endpoints and the unified `/waves/v1/tts` route, across REST/SSE (OpenAPI) and WebSocket (AsyncAPI).
* Fully backwards compatible - the parameter is optional and omitting it preserves today's behaviour exactly.

## Lightning v3.1 Pro goes multilingual - 27 new languages, 149 new voices

> The Pro pool now speaks 29 languages. 149 new premium voices across 9 Indian, 8 Asian & Middle Eastern, and 10 European languages are live on lightning_v3.1_pro.

**Lightning v3.1 Pro** now supports **27 additional languages** beyond English + Hindi, with **149 new premium voices** live on prod:

* **9 Indian languages** (67 voices) - Marathi, Tamil, Malayalam, Telugu, Kannada, Punjabi, Bengali, Odia, Gujarati
* **8 Asian & Middle Eastern languages** (21 voices) - Arabic, Chinese (Mandarin), Indonesian, Japanese, Korean, Malay, Turkish, Vietnamese
* **10 European languages** (61 voices) - German, Spanish, French, Italian, Portuguese (Brazilian + European), Russian, Greek, Finnish, Norwegian, Polish

Pass the ISO 639-1 code in the `language` body parameter and pair it with a Pro voice from that language:

```bash
curl -X POST "https://api.smallest.ai/waves/v1/tts" \
  -H "Authorization: Bearer $SMALLEST_API_KEY" \
  -H "Accept: audio/wav" \
  -d '{"text": "வணக்கம்!", "voice_id": "malar", "model": "lightning_v3.1_pro", "language": "ta", "sample_rate": 24000}'
```

**What's in this release:**

* Each language has its own dedicated Pro voices - see the [Lightning v3.1 Pro voice catalog](/model-cards/text-to-speech/lightning-v-3-1-pro#voice-catalog) for the full per-language list with voice IDs and genders.
* `language: en`, `language: hi`, and the omit-`language` default (`en + hi`) behave exactly as before.
* The full updated voice list is available via the [Get Voices endpoint](/api-reference/models/text-to-speech/get-waves-voices).

**Migration:** no action - additive change, existing voices and language behaviour unchanged.

## Lightning TTS WebSocket - documented `?timeout=N` connection-timeout knob

> The Stream Speech (WebSocket) API ref now documents the `?timeout=N` query parameter for overriding the 60-second default idle-close behavior on `WSS /waves/v1/tts/live`. Wire behavior unchanged.

The Lightning TTS WebSocket ([Stream Speech (WebSocket)](/api-reference/models/text-to-speech/tts)) API ref now documents the `?timeout=N` query parameter that has been available on the endpoint all along but wasn't surfaced in the developer-facing docs.

## What changed

Purely a docs change - no protocol or wire-level change. The new section in the API ref clarifies:

* The default idle timeout on `WSS /waves/v1/tts/live` is **60 seconds**, not the 20 seconds the legacy "WebSocket Support for TTS" page used to claim.
* Override the value with `?timeout=N` on the connection URL (positive integer seconds, e.g. `wss://api.smallest.ai/waves/v1/tts/live?timeout=120`).
* Custom values are honored verbatim, including small ones (`?timeout=5` closes after 5 seconds of silence) and large ones (verified up to `?timeout=999`).
* The timeout resets on every message you send (binary audio in, JSON control in), so keep-alive traffic restarts the clock.

## Why this matters

Voice agents with long human-thinking windows, agentic pipelines that round-trip to an LLM between TTS bursts, and any workflow with extended natural pauses now have a documented way to keep the WebSocket open past 60 seconds without resorting to dummy keep-alive frames.

## Cleanup

The standalone `/api-reference/models/web-socket` page (which previously held this info, with a stale 20-second default and a reference to the deprecated `/waves/v1/lightning-v3.1/get_speech/stream` URL) has been removed. A redirect from the old URL points at the new home.

## Lightning v3.1 - `auto` language removed from docs and spec enums

> The language: 'auto' value is no longer documented or listed in the Lightning v3.1 or Lightning v3.1 Pro spec enums.

The `language: "auto"` value is no longer documented or listed in the Lightning v3.1 or Lightning v3.1 Pro spec enums. Pass an explicit language code that matches the voice instead.

**Why:** code-switching guidance lives on the voice, not the request. Each voice in the catalog has a `tags.language` set returned by `GET /waves/v1/lightning-v3.1/get_voices`; pass a language the voice was trained on to get the pronunciation you expect. The `auto` value never actually drove language detection at the model level - it was a permissive enum value that resolved to the voice's default behavior - so removing it from the contract is the honest move.

**What changed:**

* Lightning v3.1 + Pro OpenAPI/AsyncAPI `language` enum no longer lists `auto`.
* Default value flipped from `"auto"` to `"en"` across the four specs (tts-openapi, lightning-v3.1-openapi, tts-ws, lightning-v3.1-ws).
* "Automatic Language Detection & Code-Switching" Tip removed from the v3.1 model card; the Auto-detect table row removed from Supported Languages.
* Pro model card: Languages row simplified to `English (en), Hindi (hi)`; the code-switching cell now points to `tags.language` rather than `auto`.
* All guides (quickstart, how-to-tts, stream-tts, overview, eval script) drop `auto` from their `language` parameter rows.
* Voice-cloning spec prose updated: the same recommendation (match the reference audio's language to the TTS request's `language`) stands on its own.

**Migration:** if you were sending `language: "auto"`, replace it with the language code that matches your voice (`en`, `hi`, `ta`, etc. - see `tags.language` on the voice via `GET /waves/v1/lightning-v3.1/get_voices`). Sending `auto` was not driving language detection in the first place; switching to an explicit code makes the output predictable.

## Lightning v3.1 - per-word timestamps on WebSocket streaming

> Opt in with `word_timestamps` on a WebSocket request to receive interleaved `word_timestamp` frames with per-word `{id, word, start, end}` timing - verbatim from input text, supported on Lightning v3.1 and v3.1 Pro base-queue English + Hindi voices.

Lightning v3.1 now exposes per-word timing events to WebSocket clients. Opt in with one flag - useful for captioning UIs, karaoke-style word highlighting, avatar lip-sync, and word-level analytics.

## What changed

Two changes to a WebSocket request: add `word_timestamps: true` and handle the new `status: "word_timestamp"` frame.

```js
ws.send(JSON.stringify({
  text: "I bought 3 cats for $100 on Dec 25th",
  voice_id: "meher",
  model: "lightning_v3.1_pro",
  sample_rate: 44100,
  output_format: "pcm",
  word_timestamps: true,        // ← ADDED
}));

ws.onmessage = (event) => {
  const msg = JSON.parse(event.data);
  switch (msg.status) {
    case "chunk":
      audioPlayer.push(Buffer.from(msg.data.audio, 'base64'));
      break;
    case "word_timestamp":       // ← NEW CASE
      const { id, word, start, end } = msg.data;
      captionTrack.push({ id, word, startSec: start, endSec: end });
      break;
    case "complete":
      audioPlayer.end();
      break;
  }
};
```

`word` is the exact substring from the input text - un-normalized. `"$100"` stays `"$100"`, `"25th"` stays `"25th"`, `"3"` stays `"3"`. Non-Latin scripts come back verbatim (e.g., Devanagari for Hindi).

`start` and `end` are floats in seconds, relative to the start of the audio stream. Frames interleave with `chunk` in audio-time order, then a single `complete` terminates the session.

## Where it works

| Surface                                                                        | Word timestamps                     |
| ------------------------------------------------------------------------------ | ----------------------------------- |
| `WSS /waves/v1/tts/live` (unified)                                             | ✅                                   |
| `WSS /waves/v1/lightning-v3.1/get_speech/stream` (legacy, retiring 2026-07-14) | ✅                                   |
| `POST /waves/v1/tts` (sync HTTP)                                               | ❌ - flag accepted, silently ignored |
| `POST /waves/v1/tts/live` (HTTP SSE)                                           | ❌ - same                            |

## Voice + language support

| Language                                      | Voice family                                                                  | Word events |
| --------------------------------------------- | ----------------------------------------------------------------------------- | ----------- |
| English (`en`)                                | Base-queue voices - `meher`, `devansh`, `kartik`, `maithili`, `liam`, `avery` | ✅           |
| Hindi (`hi`)                                  | Base-queue voices (same list)                                                 | ✅           |
| Marathi / Bengali / Gujarati / Punjabi / Odia | north-Indic family                                                            | ❌           |
| Tamil / Telugu / Kannada / Malayalam          | south-Indic family                                                            | ❌           |

For unsupported voice families the flag is accepted - audio works normally, but no `word_timestamp` frames are emitted. Detect this client-side by counting received word events after `complete` arrives.

## Backward compatibility

`word_timestamps` defaults to `false`. Clients that don't set the flag see no behavior change - same audio chunks, same completion frame, no new event type to handle. Purely opt-in.

**Migration:** none - pure addition. Existing integrations keep working untouched.

→ [Word-level timestamps on the Lightning v3.1 model card](/model-cards/text-to-speech/lightning-v-3-1#word-level-timestamps) - full wire spec, JS example, support matrix.

## Lightning v2 and Lightning Large endpoints retired - 410 Gone with v3.1 migration pointer

> The legacy `/lightning-v2/*` and `/lightning-large/*` endpoints now return `410 MODEL_DEPRECATED` instead of hanging. Voice cloning defaults to v3.1.

The underlying inference pools for **Lightning v2** and **Lightning Large** have been retired. Calls to these endpoints now return a fast `410 Gone` with a migration pointer to Lightning v3.1.

| Endpoint                                      | Before                       | After                                                                      |
| --------------------------------------------- | ---------------------------- | -------------------------------------------------------------------------- |
| `POST /waves/v1/lightning-v2/*`               | 5xx after timeout            | **410 `MODEL_DEPRECATED`** → use `lightning-v3.1`                          |
| `POST /waves/v1/lightning-large/*`            | 5xx after timeout            | **410 `MODEL_DEPRECATED`** → use `lightning-v3.1`                          |
| `POST /waves/v1/prof-voice-cloning/*`         | 5xx after timeout            | **410 `MODEL_DEPRECATED`** → use `/waves/v1/voice-cloning` (v3.1)          |
| `POST /waves/v1/voice-cloning` (`model=v2`)   | served via lightning-large   | **400** - Voice cloning for lightning-v2 is deprecated; use lightning-v3.1 |
| `POST /waves/v1/voice-cloning` (no `model`)   | defaulted to v2 → dead queue | **defaults to v3.1**, served                                               |
| `POST /waves/v1/voice-cloning` (`model=v3.1`) | served                       | served (no change)                                                         |

**Response shape on the deprecated endpoints:**

```json
{
  "status": "error",
  "error_code": "MODEL_DEPRECATED",
  "message": "This model is retired. Please migrate to lightning-v3.1 via /waves/v1/lightning-v3.1/get_speech.",
  "recommended_endpoint": "/waves/v1/lightning-v3.1/get_speech"
}
```

**Migration:** replace any `lightning-v2` or `lightning-large` calls with the equivalent `lightning-v3.1` endpoint. Voice cloning with no `model` parameter now routes to v3.1 automatically - no client change needed for that case.

**Not affected** (so v3.1 voice cloning keeps working): `lightning-v3.1/get_speech`, `voice-cloning` clone-creation with v3.1, and the unified `/tts` and `/tts/live` routes.

## Lightning v2 and Lightning Large endpoints retired - 410 Gone with v3.1 migration pointer

> The legacy `/lightning-v2/*` and `/lightning-large/*` endpoints now return `410 MODEL_DEPRECATED` instead of hanging. Voice cloning defaults to v3.1.

The underlying inference pools for **Lightning v2** and **Lightning Large** have been retired. Calls to these endpoints now return a fast `410 Gone` with a migration pointer to Lightning v3.1.

| Endpoint                                      | Before                       | After                                                                      |
| --------------------------------------------- | ---------------------------- | -------------------------------------------------------------------------- |
| `POST /waves/v1/lightning-v2/*`               | 5xx after timeout            | **410 `MODEL_DEPRECATED`** → use `lightning-v3.1`                          |
| `POST /waves/v1/lightning-large/*`            | 5xx after timeout            | **410 `MODEL_DEPRECATED`** → use `lightning-v3.1`                          |
| `POST /waves/v1/prof-voice-cloning/*`         | 5xx after timeout            | **410 `MODEL_DEPRECATED`** → use `/waves/v1/voice-cloning` (v3.1)          |
| `POST /waves/v1/voice-cloning` (`model=v2`)   | served via lightning-large   | **400** - Voice cloning for lightning-v2 is deprecated; use lightning-v3.1 |
| `POST /waves/v1/voice-cloning` (no `model`)   | defaulted to v2 → dead queue | **defaults to v3.1**, served                                               |
| `POST /waves/v1/voice-cloning` (`model=v3.1`) | served                       | served (no change)                                                         |

**Response shape on the deprecated endpoints:**

```json
{
  "status": "error",
  "error_code": "MODEL_DEPRECATED",
  "message": "This model is retired. Please migrate to lightning-v3.1 via /waves/v1/lightning-v3.1/get_speech.",
  "recommended_endpoint": "/waves/v1/lightning-v3.1/get_speech"
}
```

**Migration:** replace any `lightning-v2` or `lightning-large` calls with the equivalent `lightning-v3.1` endpoint. Voice cloning with no `model` parameter now routes to v3.1 automatically - no client change needed for that case.

**Not affected** (so v3.1 voice cloning keeps working): `lightning-v3.1/get_speech`, `voice-cloning` clone-creation with v3.1, and the unified `/tts` and `/tts/live` routes.

## Lightning v3.1 Pro - premium voice catalog across American, British, and Indian accents

> Lightning v3.1 now has a Pro tier with 39 curated voices across American, British, and Indian accents (both Male and Female).

Lightning v3.1 now has a **Pro** tier with **39 curated voices** across American, British, and Indian accents (both Male and Female). The Pro pool runs on dedicated inference capacity, delivering the same TTFB as standard Lightning v3.1.

**What's in the catalog:**

* **Indian - Female (8)**: Rhea, Zariya, Kareena, Mishka, Inaaya, Saira, Meher, Aarini
* **Indian - Male (5)**: Aviraj, Vyom, Zoravar, Reyansh, Ahan
* **British - Female (6)**: Cressida, Elowen, Ottilie, Seraphina, Tabitha, Arabella
* **British - Male (7)**: Benedict, Cormac, Everett, Finley, Rupert, Winston, Caspian
* **American - Female (7)**: Willow, Autumn, Skylar, Savannah, Kennedy, Reagan, Sierra
* **American - Male (6)**: Maverick, Brooks, Hunter, Colton, Wesley, Asher

**Languages supported:** Indian voices speak English and Hindi (with native code-switching when `language="auto"`). British and American voices speak English. See per-voice `tags.language` via `GET /waves/v1/lightning-v3.1/get_voices`.

**How to use it:**

* **Atoms voice agents:** open the agent's voice picker and select the new **Pro** filter chip, then pick any Pro voice. Atoms transparently routes to the Pro pool - no other configuration needed.
* **API (direct):** use the unified `POST /waves/v1/tts` (sync), `POST /waves/v1/tts/live` (SSE), or `WSS /waves/v1/tts/live` (WebSocket) endpoints and pass `"model": "lightning_v3.1_pro"` in the request body alongside the chosen `voice_id`. The legacy `/waves/v1/lightning-v3.1/*` routes also accept the `model` field for backwards-compatible Pro opt-in.

**Voice cloning:** not available on Lightning v3.1 Pro. Voice clones continue to use Lightning v3.1 (standard) and the existing voice-cloning flow. There is no migration required.

For the full catalog, integration examples, and a Python WebSocket sample, see the [Lightning v3.1 Pro model card](/model-cards/text-to-speech/lightning-v-3-1-pro).

## Indic voices now produce clean audio regardless of the `language` field

> Lightning v3.1 used to pick an inference pool from the request's language field, which meant Indic voices (Aadya, Yuvan, Samarth, Nilesh, Arnab, Niharika.

Lightning v3.1 used to pick an inference pool from the request's `language` field, which meant Indic voices (Aadya, Yuvan, Samarth, Nilesh, Arnab, Niharika, Gargi, and other voices whose latents are trained on north\_indic / south\_indic encoders) could be served from the wrong pool when called with `language=en` or `language=hi` - producing distorted or unintelligible audio.

Routing now derives from the **voice itself**, not the request language:

* Voices tagged `odia` / `bengali` / `punjabi` / `gujarati` / `marathi` route to the **north\_indic** inference pool
* Voices tagged `kannada` / `malayalam` / `telugu` / `tamil` route to the **south\_indic** pool
* All other voices continue to route by language as before

**No code change needed.** If you had previously worked around this by hard-coding `language` to match the voice family, you can remove that workaround - the platform now picks the correct pool automatically.

**Voice clones are unaffected** - clones bypass this lookup since they aren't in the public voice catalog.

## Lightning v3.1 - language list corrected to 12 (voice catalog source of truth)

> Plus auto for automatic language detection and code-switching across the above set.

**Correction.** Earlier in the week we expanded the Lightning v3.1 documented language list to 22 codes plus `auto`, sourced from the server-side `lightningV3_1Schema` enum in waves-platform. Live testing showed those 22 codes are *accepted by the schema* but only **12 of them have voices in the catalog** - the other 10 (`de`, `fr`, `it`, `pl`, `nl`, `ru`, `sv`, `pt`, `ar`, `he`) silently fall back to the voice's default language when called. We were lying to users.

**Source of truth is now the voice catalog**, not the schema enum. The actually-supported set is:

| Code | Language  | Voices |
| ---- | --------- | -----: |
| `en` | English   |    176 |
| `hi` | Hindi     |    115 |
| `ta` | Tamil     |     13 |
| `es` | Spanish   |     11 |
| `kn` | Kannada   |     10 |
| `mr` | Marathi   |      9 |
| `te` | Telugu    |      8 |
| `or` | Odia      |      8 |
| `pa` | Punjabi   |      8 |
| `ml` | Malayalam |      6 |
| `gu` | Gujarati  |      5 |
| `bn` | Bengali   |      4 |

Plus `auto` for automatic language detection and code-switching across the above set. **217 voices total.**

**What changed:**

* Lightning v3.1 OpenAPI + AsyncAPI `language` enum narrowed to 12 codes + `auto`.
* Model card, getting-started/models, text-to-speech overview, api-references/lightning-v3.1, integrations (LiveKit, Vercel AI SDK, JellyPod) all updated.
* Removed the bogus `Beta` rows for German, French, Italian, Polish, Dutch, Russian, Swedish, Portuguese, Arabic, Hebrew.
* Voice count corrected from "169 voices total" to 217.

**Pulse STT (separate model) is unaffected** - Pulse genuinely supports its full European + Indic + Asian language set via the `multi-eu`, `multi-indic`, `multi-asian` regional aggregators.

**Reproducible verification.** A live probe (`scripts/spec-live-tests/spec_enum_vs_voice_catalog.py`) now compares the spec's language enum against the live `GET /lightning-v3.1/get_voices` response and fails CI if they drift. This catches the schema-vs-reality gap on every spec PR going forward.

## Legacy Lightning STT/TTS API reference orphans removed from docs

> Several legacy API reference MDX files that had already been unlinked from the v4 API reference navigation have been removed from the docs.

Several legacy API reference MDX files that had already been unlinked from the v4 API reference navigation have been removed from the docs.

**STT (Speech-to-Text):**

* The legacy `Lightning (Pre-Recorded)` HTTP reference (`POST /waves/v1/lightning/get_text`) and its OpenAPI spec (`fern/apis/waves/openapi/asr-openapi.yaml`). The current STT pre-recorded surface is **Pulse** (`POST /waves/v1/pulse/get_text`) - see [Pulse pre-recorded reference](/api-reference/models/speech-to-text/transcribe). The legacy MDX page was a verbatim copy of the Pulse one with the URL substituted, so no functionality is lost.
* (PR #110 separately removes the matching `Lightning ASR WebSocket` reference and its AsyncAPI spec.)

**TTS (Text-to-Speech):**

* The legacy `lightning-large` HTTP TTS, SSE, and WebSocket reference pages (`lightning-large.mdx`, `lightning-large-stream.mdx`, `lightning-large-ws.mdx`). These were already unlinked from v4 nav. The current TTS surface is `lightning-v3.1` - the migration prose under [Voice Cloning: Instant Clone API](/models/voice-cloning/instant-clone-api) already documents the cutover.

**Voice cloning impact:** none. The current voice-cloning flow (`POST /waves/v1/voice-cloning`) is unchanged. The deprecated `lightning-large` endpoints that *are* still in use (`add_voice`, `get_cloned_voices`, `DELETE /waves/v1/lightning-large`) remain in the API reference under **Voice Cloning** with their existing `(Deprecated)` labels.

If your code calls `https://api.smallest.ai/waves/v1/lightning/get_text` for STT or `https://api.smallest.ai/waves/v1/lightning-large/get_speech` for TTS, switch to Pulse and Lightning v3.1 respectively.

→ [Pulse pre-recorded reference](/api-reference/models/speech-to-text/transcribe)
→ [TTS API reference](/api-reference/models/text-to-speech/synthesize-speech)

## Legacy Lightning STT/TTS API reference orphans removed from docs

> Several legacy API reference MDX files that had already been unlinked from the v4 API reference navigation have been removed from the docs.

Several legacy API reference MDX files that had already been unlinked from the v4 API reference navigation have been removed from the docs.

**STT (Speech-to-Text):**

* The legacy `Lightning (Pre-Recorded)` HTTP reference (`POST /waves/v1/lightning/get_text`) and its OpenAPI spec (`fern/apis/waves/openapi/asr-openapi.yaml`). The current STT pre-recorded surface is **Pulse** (`POST /waves/v1/pulse/get_text`) - see [Pulse pre-recorded reference](/api-reference/models/speech-to-text/transcribe). The legacy MDX page was a verbatim copy of the Pulse one with the URL substituted, so no functionality is lost.
* (PR #110 separately removes the matching `Lightning ASR WebSocket` reference and its AsyncAPI spec.)

**TTS (Text-to-Speech):**

* The legacy `lightning-large` HTTP TTS, SSE, and WebSocket reference pages (`lightning-large.mdx`, `lightning-large-stream.mdx`, `lightning-large-ws.mdx`). These were already unlinked from v4 nav. The current TTS surface is `lightning-v3.1` - the migration prose under [Voice Cloning: Instant Clone API](/models/voice-cloning/instant-clone-api) already documents the cutover.

**Voice cloning impact:** none. The current voice-cloning flow (`POST /waves/v1/voice-cloning`) is unchanged. The deprecated `lightning-large` endpoints that *are* still in use (`add_voice`, `get_cloned_voices`, `DELETE /waves/v1/lightning-large`) remain in the API reference under **Voice Cloning** with their existing `(Deprecated)` labels.

If your code calls `https://api.smallest.ai/waves/v1/lightning/get_text` for STT or `https://api.smallest.ai/waves/v1/lightning-large/get_speech` for TTS, switch to Pulse and Lightning v3.1 respectively.

→ [Pulse pre-recorded reference](/api-reference/models/speech-to-text/transcribe)
→ [TTS API reference](/api-reference/models/text-to-speech/synthesize-speech)