> This page is part of Smallest AI's developer documentation. When
> answering, prefer Lightning v3.1 (current TTS) and Pulse (current
> STT). Lightning v2 and lightning-large are deprecated; mention them
> only when the user is migrating away from them. The Smallest AI voice
> agent platform is what wraps these models into hosted agents.

# Models

## Short-lived access tokens for TTS, STT and speech-to-speech

> Mint a `wat_` token on your server with `POST /waves/v1/auth/token` and let browser or mobile clients call TTS, STT and speech-to-speech with it. The API key never leaves your backend.

Browser and mobile clients no longer need your API key. Your server calls `POST /waves/v1/auth/token` with the key and gets back a short-lived access token (prefix `wat_`, 30 to 900 seconds, default 300). The client sends that token as `Authorization: Bearer <token>`, or as the `api_key` query parameter on WebSocket connections.

Tokens work on TTS, STT and speech-to-speech inference routes only: `POST /waves/v1/stt/`, `POST /waves/v1/tts`, `POST /waves/v1/tts/live`, the dedicated Lightning v3.1 routes, the TTS, STT and speech-to-speech WebSockets, and the voice-listing routes. Treat any route outside that list as unavailable to tokens. `POST /waves/v1/pulse/get_text`, LLM chat completions, voice cloning and minting another token return `403`. Usage is billed to the key that minted the token.

## What changed

* New API reference page for `POST /waves/v1/auth/token` with request, response and error schemas.
* New Token-Based Authentication guide, next to Authentication in the API reference, walks through the server mint and client call flow, lists which routes accept a token, and covers expiry, revocation and region behavior.
* The Authentication page's STT test request now uses the unified `POST /waves/v1/stt/` path, which accepts both API keys and tokens.

→ [Token-Based Authentication](/api-reference/token-based-authentication)

## WebSocket keep-alive frame is now documented

> Send `{"type":"ping"}` on any Waves WebSocket to hold an idle connection open and get a `pong` back. Available since August; the docs and API reference now cover it, along with the real session limits.

Every Waves WebSocket - Pulse STT and Lightning TTS - accepts an application-level keep-alive frame. Send `{"type": "ping"}` and the server resets the inactivity timer and replies `{"type": "pong"}`. `keepalive`, `keep_alive`, and `keep-alive` work as aliases.

The frame is handled at the API edge: it never reaches the model, does not affect the transcript or the synthesized audio, and carries no audio, so it is not billed.

## What changed

Documentation only - the wire behavior has been live since August. The new material covers:

* The keep-alive frame and its `pong` acknowledgement, on both the STT and TTS WebSocket references.
* The three limits that apply to a Pulse STT session: 20-minute connection inactivity (raisable to 3 hours with `?timeout=N`, and reset by keep-alives), a 20-minute model-session idle window that keep-alives do **not** extend, and a 5-hour absolute session lifetime.
* A correction on the TTS WebSocket reference: `?timeout=N` is capped at **180 seconds**, not honored verbatim at any value as previously stated.

→ [Keep-Alive](/models/documentation/speech-to-text-pulse/features/keep-alive)

## Docs: fix five stale internal links after slug renames

> Five internal links in the Models docs now point at the correct pages: - Electron model card API-reference link (electron-chat-completions) - LLM overview + q

Five internal links in the Models docs now point at the correct pages:

* Electron model card API-reference link (`electron-chat-completions`)
* LLM overview + quickstart `migrate-from-openai` link
* Pulse STT realtime features `age-and-gender-detection` link
* Lightning v3.1 Pro model card `voice-cloning/how-to-vc-api` link

Also updates the `fern/docs.yml` redirects table so old URLs keep resolving. No product behavior change.

## Speech-to-Speech: Hydra V1.1 released, `?model=hydra` deprecated

> Hydra V1.1 is the current release of the realtime speech-to-speech model.

Hydra V1.1 is the current release of the realtime speech-to-speech model. New integrations should connect with `?model=hydra-v1.1`:

```
wss://api.smallest.ai/waves/v1/s2s?model=hydra-v1.1&api_key=<SMALLEST_API_KEY>
```

The session protocol (event catalog, `session.configure` shape, tool calling, interruption handling) is identical to the original `hydra`, so switching is a one-parameter change on the query string.

**`?model=hydra` is deprecated.** Existing sessions still open, but the server emits a `warning` frame with `code: "model_deprecated"` immediately before `session.created`. Migrate to `?model=hydra-v1.1`. See [Deprecation Notices](/api-reference/deprecations/models).

**Voice rosters** are documented per version on the [Hydra model card](/model-cards/speech-to-speech/hydra). An unknown voice on `session.configure` is rejected with an `error` frame (`code: "invalid_request_error"`) rather than silently defaulted, so validate client-side.

**`hydra-v1.0` is also served** on the same endpoint; it is not deprecated, and the server names it as the migration target when a client connects with the deprecated `?model=hydra`.

Full voice + version reference on the [Hydra model card](/model-cards/speech-to-speech/hydra). Overview + docs pages ([overview](/models/speech-to-speech/overview), [WebSocket connection](/models/speech-to-speech/web-socket-connection), [Audio I/O](/models/speech-to-speech/audio-i-o), [Managing sessions](/models/speech-to-speech/managing-sessions), [Tool calling](/models/speech-to-speech/tool-calling)) point at `?model=hydra-v1.1` in every URL example and use a V1.1 voice in the sample payloads. The `voice` field in **Managing sessions** points to the model card, which is the single source of truth for per-version voice tables, and documents the rejection of unknown voices.

## Speech-to-Speech: Hydra V1.1 released, `?model=hydra` deprecated

> Hydra V1.1 is the current release of the realtime speech-to-speech model.

Hydra V1.1 is the current release of the realtime speech-to-speech model. New integrations should connect with `?model=hydra-v1.1`:

```
wss://api.smallest.ai/waves/v1/s2s?model=hydra-v1.1&api_key=<SMALLEST_API_KEY>
```

The session protocol (event catalog, `session.configure` shape, tool calling, interruption handling) is identical to the original `hydra`, so switching is a one-parameter change on the query string.

**`?model=hydra` is deprecated.** Existing sessions still open, but the server emits a `warning` frame with `code: "model_deprecated"` immediately before `session.created`. Migrate to `?model=hydra-v1.1`. See [Deprecation Notices](/api-reference/deprecations/models).

**Voice rosters** are documented per version on the [Hydra model card](/model-cards/speech-to-speech/hydra). An unknown voice on `session.configure` is rejected with an `error` frame (`code: "invalid_request_error"`) rather than silently defaulted, so validate client-side.

**`hydra-v1.0` is also served** on the same endpoint; it is not deprecated, and the server names it as the migration target when a client connects with the deprecated `?model=hydra`.

Full voice + version reference on the [Hydra model card](/model-cards/speech-to-speech/hydra). Overview + docs pages ([overview](/models/speech-to-speech/overview), [WebSocket connection](/models/speech-to-speech/web-socket-connection), [Audio I/O](/models/speech-to-speech/audio-i-o), [Managing sessions](/models/speech-to-speech/managing-sessions), [Tool calling](/models/speech-to-speech/tool-calling)) point at `?model=hydra-v1.1` in every URL example and use a V1.1 voice in the sample payloads. The `voice` field in **Managing sessions** points to the model card, which is the single source of truth for per-version voice tables, and documents the rejection of unknown voices.

## STT: new language `hi-dev` (Devanagari-only Hindi)

> Pulse STT streaming supports a new Hindi language code, hi-dev, that transcribes Hindi audio using a pure-Hindi model and writes everything in Devanagari.

Pulse STT streaming supports a new Hindi language code, `hi-dev`, that transcribes Hindi audio using a pure-Hindi model and writes everything in Devanagari.

The existing `hi` code stays as the bilingual model: it transcribes Hindi speech but renders borrowed English words in Roman script, producing mixed-script output like `मेरी email आईडी पर invoice भेज दीजिए`.

`hi-dev` uses the pure-Hindi model to keep the output script-consistent: `मेरी ईमेल आईडी पर इनवॉइस भेज दीजिए`.

Pick `hi-dev` when the downstream consumer needs Devanagari-only text (Devanagari-only UIs, government or regulated formatting, Hindi NLP pipelines that break on Latin characters). Stay on `hi` for customer-support and tech-support transcripts where recognisable English terms in Roman script are preferable.

**Numbers.** Pass `itn_normalize=true` to convert numbers spoken in Hindi to digits (`पच्चीस हज़ार रुपए` becomes `₹25000`). Numbers spoken in English are written as Devanagari words and are not converted.

Endpoint: `wss://api.smallest.ai/waves/v1/stt/live?language=hi-dev`. Same auth, message shape, and parameters as every other language.

Streaming only; not available on `POST /waves/v1/stt/`. India region only; requests routed to `wss://api.us.smallest.ai/...` are rejected with `LANGUAGE_NOT_ENABLED_IN_REGION`.

## Speech to Text: VAD events on the live WebSocket API reference

> The Pulse STT live WebSocket now models the acoustic voice-activity events in the API reference: speech_started and speech_ended, emitted alongside.

The Pulse STT live WebSocket now models the acoustic voice-activity events in the API reference: `speech_started` and `speech_ended`, emitted alongside `transcription` messages when the connection sets `vad_events=true`. See [VAD events](/models/speech-to-text/features/vad-events) for the payloads and usage.

The events have been in production; this documents them on the WebSocket reference.

## Waves: analytics endpoints on the API reference

> Nine previously undocumented Waves analytics endpoints now render on the API reference under a new Analytics section:

Nine previously undocumented Waves analytics endpoints now render on the API reference under a new **Analytics** section:

* `GET /waves/v1/analytics/asr/logs`: paginated STT request log
* `DELETE /waves/v1/analytics/asr/history/{request_id}`: soft-delete a single STT entry
* `GET /waves/v1/analytics/asr/usage/timeseries`: STT request count over time
* `GET /waves/v1/analytics/tts/logs`: paginated TTS request log
* `GET /waves/v1/analytics/tts/usage/timeseries`: TTS request count over time
* `GET /waves/v1/analytics/tts/usage/credits/timeseries`: TTS credit spend over time
* `GET /waves/v1/analytics/tts/concurrency/timeseries`: TTS peak concurrency over time
* `GET /waves/v1/analytics/tts/ws-connections/timeseries`: open TTS WebSocket count over time
* `GET /waves/v1/analytics/webhooks/logs`: webhook delivery log (e.g. `asr.completed` callbacks)

Timeseries endpoints accept an ISO 8601 datetime range (`from`, `to`) and an optional `granularity` bucket. Log endpoints are paginated with `page` and `pageSize`. All routes are scoped by API key to the caller's organization.

## Waves: analytics endpoints on the API reference

> Nine previously undocumented Waves analytics endpoints now render on the API reference under a new Analytics section:

Nine previously undocumented Waves analytics endpoints now render on the API reference under a new **Analytics** section:

* `GET /waves/v1/analytics/asr/logs`: paginated STT request log
* `DELETE /waves/v1/analytics/asr/history/{request_id}`: soft-delete a single STT entry
* `GET /waves/v1/analytics/asr/usage/timeseries`: STT request count over time
* `GET /waves/v1/analytics/tts/logs`: paginated TTS request log
* `GET /waves/v1/analytics/tts/usage/timeseries`: TTS request count over time
* `GET /waves/v1/analytics/tts/usage/credits/timeseries`: TTS credit spend over time
* `GET /waves/v1/analytics/tts/concurrency/timeseries`: TTS peak concurrency over time
* `GET /waves/v1/analytics/tts/ws-connections/timeseries`: open TTS WebSocket count over time
* `GET /waves/v1/analytics/webhooks/logs`: webhook delivery log (e.g. `asr.completed` callbacks)

Timeseries endpoints accept an ISO 8601 datetime range (`from`, `to`) and an optional `granularity` bucket. Log endpoints are paginated with `page` and `pageSize`. All routes are scoped by API key to the caller's organization.

## STT API-ref: response examples now match the real API

> Follow-up to yesterday's word_timestamps clarification.

Follow-up to yesterday's `word_timestamps` clarification. That PR fixed the description but the response schema itself still drifted from what the live API returns. Postman side-by-side surfaced it.

Corrected on both endpoint schemas (`/waves/v1/stt/` and legacy `/waves/v1/pulse/get_text`):

* **`speaker` is an integer, not a string.** Real responses carry `speaker: 0` (zero-indexed), not `speaker: "speaker_0"`.
* **`speaker_confidence` field added** to word entries. Present alongside `speaker` when `diarize=true` is set on Pulse.
* **Pulse metadata is `{duration, fileSize}` only.** The old example claimed `language`, `request_id`, `processing_time_ms`, `rtfx`, and `num_chunks` on a Pulse response; those are Pulse Pro only.
* **Pulse Pro includes `totalBytes`, `request_id`, and `language`** at the top level and richer `metadata` (`processing_time_ms`, `rtfx`, `num_chunks`). Also does not diarize, so `words[]` entries omit `speaker` and `speaker_confidence`, and there is no `utterances[]` field.
* **`duration` is seconds, not minutes.** Description on the legacy Pulse endpoint schema said minutes; corrected.
* **Response examples split by scenario.** The unified endpoint now shows three named examples: `pulse-default` (empty arrays), `pulse-full` (word\_timestamps + diarize), and `pulse-pro`. The legacy endpoint shows two named examples: default and with word\_timestamps + diarize.

Live-verified against `api.smallest.ai` on a WAV generated via TTS: default, word\_timestamps + diarize, Pulse Pro, and webhook async paths all match the updated schemas.

## STT API: clarify that words[] and utterances[] need word_timestamps=true

> The STT API-ref example response showed populated words[] and utterances[] arrays, but a default request from the docs explorer or Postman returns both.

The STT API-ref example response showed populated `words[]` and `utterances[]` arrays, but a default request from the docs explorer or Postman returns both arrays empty (the `transcription` string is fine).

Both arrays are gated on the `word_timestamps=true` query parameter. The same flag turns on both. Without it, the API still transcribes but does not attach per-word timings or sentence-level segments.

Fixed on both endpoint schemas (`/waves/v1/stt/` and legacy `/waves/v1/pulse/get_text`):

* Response schema `words` and `utterances` descriptions now state the gate.
* Example responses now carry a YAML comment naming the request shape they correspond to (word\_timestamps=true, plus diarize=true where the example includes `speaker`).
* `word_timestamps` query-parameter description now spells out default-off behavior and links to the Word Timestamps feature page.

Live-verified: same wav sample against both endpoints returned `transcription` populated with empty `words[]` / `utterances[]` when `word_timestamps` was omitted, and 12 words + 2 utterances when the flag was set.

## STT API: clarify that words[] and utterances[] need word_timestamps=true

> The STT API-ref example response showed populated words[] and utterances[] arrays, but a default request from the docs explorer or Postman returns both.

The STT API-ref example response showed populated `words[]` and `utterances[]` arrays, but a default request from the docs explorer or Postman returns both arrays empty (the `transcription` string is fine).

Both arrays are gated on the `word_timestamps=true` query parameter. The same flag turns on both. Without it, the API still transcribes but does not attach per-word timings or sentence-level segments.

Fixed on both endpoint schemas (`/waves/v1/stt/` and legacy `/waves/v1/pulse/get_text`):

* Response schema `words` and `utterances` descriptions now state the gate.
* Example responses now carry a YAML comment naming the request shape they correspond to (word\_timestamps=true, plus diarize=true where the example includes `speaker`).
* `word_timestamps` query-parameter description now spells out default-off behavior and links to the Word Timestamps feature page.

Live-verified: same wav sample against both endpoints returned `transcription` populated with empty `words[]` / `utterances[]` when `word_timestamps` was omitted, and 12 words + 2 utterances when the flag was set.

## New `x-expire-content` header: opt a request into content deletion (enterprise)

> Documented as a header parameter on the HTTP routes - POST /waves/v1/tts, POST /waves/v1/tts/live, POST /waves/v1/stt/ - and on the WebSocket routes.

**Enterprise plans only.** Send `x-expire-content: true` on a Text to Speech or Speech to Text request and that request's free-text content is deleted after 7 days, while the usage record is kept - billing, credits, and usage graphs are unaffected.

Documented as a header parameter on the HTTP routes - `POST /waves/v1/tts`, `POST /waves/v1/tts/live`, `POST /waves/v1/stt/` - and on the WebSocket routes `wss://.../waves/v1/tts/live` and `wss://.../waves/v1/stt/live`, where it is sent on the upgrade request.

**Nothing changes unless you send the header.** Without it, content is retained indefinitely exactly as before. The header **deletes** content - it does not mean "retain my data", and opting in is the destructive direction, so it is never applied implicitly.

What is deleted per product:

* **Speech to Text**: `transcription`. Duration, language, model, latency, request ID, and emotion/gender/age detections are kept.
* **Text to Speech**: input `text` and `normalized_text`. Credits, text length, voice ID, model, speed, and output format are kept.
* **LLM**: `messages`, `response_content`, `tool_calls`, `request_params`. Token counts, model, latency, and finish reason are kept.

**Check the `x-content-expiry` response header to confirm it took effect.** A request from a non-enterprise plan is *ignored rather than rejected* - it succeeds and content is retained - so the response tells you which happened:

* `applied` - content will be deleted after the window.
* `not-entitled` - plan does not include this; content retained.
* `unavailable` - entitlement could not be verified; content retained, safe to retry.

The header is absent when you did not request expiry. WebSocket sessions have no response headers, so the opt-in works on the upgrade but the outcome cannot be read back - use an HTTP endpoint to confirm your entitlement.

Two behaviours to plan for:

* **Deletion is approximate.** Content becomes unreadable shortly *after* the 7-day window rather than at the exact second. Do not rely on it as a hard contractual deadline without talking to us.
* **Log reads and search reflect it.** Expired content returns as an empty string from the log endpoints, and a `text` search cannot match it. The row still counts toward `totalCount`, so a result count can exceed the visible matches.

## TTS: latency measurement caveat + legacy /streaming-tts/stream artifacts retired

> Two related updates to the TTS docs.

Two related updates to the TTS docs.

**Latency measurement caveat added to the Lightning v3.1 + Pro model cards.** The published \~200 ms TTFB is measured from a client in the same AWS region as the inference pool (India `ap-south-1`, USA `us-west-2`, geo-routed). Client-to-server RTT adds on top. A naive TTFB test from a laptop far from either region typically reports 500-800 ms because RTT dominates synthesis time. Callout now spells this out next to the number so readers know to test from their production network position, not their dev machine.

**Legacy `/waves/v1/streaming-tts/stream` artifacts retired.** The `lightning-tts-ws.mdx` API-reference page and the `stream-tts-ws.yaml` AsyncAPI + overrides are removed. Both predated the unified `/waves/v1/tts/live` split and documented the SSE transport with the WebSocket envelope shape, which was contradictory. The unified endpoint has been the documented and supported path for a while; these files were dormant. Generator wiring, spec-drift allow-list, and nav-ignore lists updated accordingly.

**Canonical TTS response shapes** on `/waves/v1/tts/live`, now documented on `stream-tts.mdx` with the exact wire fields (session/request IDs on WS, `status` + `done` on SSE) plus a porting callout so a WebSocket parser doesn't silently drop SSE frames.

* WebSocket chunks: `{"session_id","request_id","status":"chunk","data":{"audio":"..."}}`.
* WebSocket terminator: `{"session_id","request_id","status":"complete"}`. **No `message`, no `done`.** Detect completion with `status == "complete"`.
* SSE chunks: `{"audio":"...","done":false,"status":"206"}`.
* SSE terminator: `{"status":"200","done":true}`. **`done` is on every frame** (`false` on chunks, `true` on terminator). Check `done == true`, not `"done" in msg`.

The two shapes are not interchangeable. A parser written for the WebSocket envelope will silently drop every SSE frame because `msg["data"]` is undefined.

## TTS language surface aligned with product-team canonical

> The unified TTS endpoints (POST /waves/v1/tts + wss://api.smallest.ai/waves/v1/tts/live) now document the full language surface the platform accepts.

The unified TTS endpoints (`POST /waves/v1/tts` + `wss://api.smallest.ai/waves/v1/tts/live`) now document the full language surface the platform accepts, matching the product-team lineup for both `lightning_v3.1` (base) and `lightning_v3.1_pro`.

**Added to the `language` enum on both HTTP and WebSocket:**

* `auto` - routes the request internally based on input text. Any English or Hindi voice can be used across all supported languages when `auto` is set; the platform handles language-appropriate routing without needing an explicit code per call. Recommended for cross-language use cases.
* `nl` (Dutch) - accepted on both `lightning_v3.1` and `lightning_v3.1_pro`.
* `sv` (Swedish) - accepted on both `lightning_v3.1` and `lightning_v3.1_pro`.

**Updated counts:**

* `lightning_v3.1` accepts **20 language codes** (was documented as 12): 10 European (English, Spanish, French, German, Italian, Dutch, Swedish, Portuguese, Polish, Russian) + 10 Indic (Hindi, Marathi, Gujarati, Punjabi, Bengali, Odia, Tamil, Telugu, Kannada, Malayalam). The trained voice catalog covers 12 of these directly; the remaining 8 route through English or Hindi voices.
* `lightning_v3.1_pro` accepts **31 language codes** (was documented as 29): the 20 above plus 3 additional European (Greek, Finnish, Norwegian) and 8 Asian & Middle Eastern (Chinese, Japanese, Korean, Indonesian, Malay, Vietnamese, Turkish, Arabic).

**No breaking changes** - existing codes still behave identically. `auto`, `nl`, and `sv` were already accepted by the platform; this update documents them.

The Lightning v3.1 and Lightning v3.1 Pro [model cards](/model-cards/text-to-speech/lightning-v-3-1) have the full per-model language breakdown and voice-count tables.

## Models docs URL migration - waves → models, version dropdown flattened

> The Models product now lives under /models/ on the docs site, replacing the older /waves/ prefix.

The Models product now lives under `/models/` on the docs site, replacing the older `/waves/` prefix. The version dropdown is also gone. There is only one active version (v4.0.0), and it is served from a flat URL - `/waves/v-4-0-0/documentation/text-to-speech-lightning/streaming` becomes `/models/documentation/text-to-speech-lightning/streaming`.

Every old URL 301-redirects to its new equivalent. Legacy v2.2.0 and v3.0.1 bookmarks (retired earlier) also land on the same flat `/models/*` paths. Bookmarks, external links, and search-indexed pages keep resolving.

API endpoints and SDK identifiers are unchanged. `POST /waves/v1/tts`, `POST /waves/v1/chat/completions`, `wss://api.smallest.ai/waves/v1/stt/live`, and the `waves` OpenAPI namespace on the SDK all still work exactly as before. Only docs URLs moved.

## Waves docs v2.2.0 and v3.0.1 retired; v4.0.0 is now the sole documented surface

> The v4.0.0 doc surface is now the only supported view of the Waves models.

The **v4.0.0** doc surface is now the only supported view of the Waves models. The version dropdown no longer offers **v3.0.1** or **v2.2.0**, and the pages under those two versions have been removed from the docs site. Old bookmarked and search-indexed URLs 301-redirect to their v4 equivalents; where a specific v3.0.1/v2.2.0 page was renamed or merged in the v4 restructure, the redirect points at the closest v4 page (voice-cloning how-tos land on `Instant Clone (API)` etc.). The API-reference tabs on the retired versions rendered a completely different OpenAPI (Lightning v1/v2/large focused), so any auto-generated API-ref URLs on those versions land on the v4 API-ref authentication page.

Nothing changes at the API layer. The OpenAPI and AsyncAPI specs remain intact; SDK generation is unaffected. This is a docs-surface cleanup only. Customers running against the legacy `POST /waves/v1/lightning-v2/*` or `POST /waves/v1/lightning-large/*` routes continue to work exactly as before, and those routes' current status is documented under the deprecated entries in the v4 changelog and via the response `x-fern-availability: deprecated` markers.

## Agno integration guide

> Added a documentation page for generating speech with Lightning v3.1 and Lightning v3.1 Pro inside Agno, the open-source Python framework for building.

Added a documentation page for generating speech with Lightning v3.1 and Lightning v3.1 Pro inside [Agno](https://github.com/agno-agi/agno), the open-source Python framework for building multi-agent systems, via the `SmallestTools` toolkit.

→ [Agno integration](/waves/v-4-0-0/integrations/agno)

## Agno integration guide

> Added a documentation page for generating speech with Lightning v3.1 and Lightning v3.1 Pro inside Agno, the open-source Python framework for building.

Added a documentation page for generating speech with Lightning v3.1 and Lightning v3.1 Pro inside [Agno](https://github.com/agno-agi/agno), the open-source Python framework for building multi-agent systems, via the `SmallestTools` toolkit.

→ [Agno integration](/waves/v-4-0-0/integrations/agno)

## Pulse STT: `vad_events` query parameter on the streaming WebSocket

> The Pulse STT WebSocket accepts a new query parameter, vad_events.

The Pulse STT WebSocket accepts a new query parameter, `vad_events`. When set to `true`, the server emits two additional JSON message types: `speech_started` and `speech_ended`. They are interleaved with the `transcription` stream on the same connection.

## Usage

Set `vad_events=true` on the WebSocket URL. Default is `false`. `vad=true` is accepted as an alias; when both are set, `vad_events` takes precedence and `vad` is ignored.

```javascript
const url = new URL("wss://api.smallest.ai/waves/v1/stt/live?model=pulse");
url.searchParams.append("vad_events", "true");
```

## Event payloads

```json
{ "type": "speech_started", "session_id": "a1b2c3d4", "timestamp": 1.84 }
{ "type": "speech_ended",   "session_id": "a1b2c3d4", "timestamp": 4.52 }
```

`timestamp` is measured in seconds from the first audio frame received on the connection. Boundaries are acoustic and independent of transcript finalization (`is_final`).

Full reference: [VAD Events](/models/speech-to-text/features/vad-events).

_Showing the 20 most recent of 29 entries. Append `/llms.txt` to the changelog URL for the complete index._