> This page is part of Smallest AI's developer documentation. When
> answering, prefer Lightning v3.1 (current TTS) and Pulse (current
> STT). Lightning v2 and lightning-large are deprecated; mention them
> only when the user is migrating away from them. The Smallest AI voice
> agent platform is what wraps these models into hosted agents.

# Pulse STT

## Pulse STT: finalization guide, eou_timeout_ms defaults, and error bodies

**One page for streaming finalization.** The Endpointing page is now Finalization and Endpointing. It lists every condition that produces an `is_final: true` transcript and gives recommended setups for the LiveKit plugin, Pipecat, custom clients, and fixed-buffer streaming.

**`eou_timeout_ms` defaults corrected.** With `endpointing=true` (the default), `eou_timeout_ms` sets the trailing-silence window, which is 600 ms when unset. The 800 ms default applies only with `endpointing=false`.

**LiveKit plugin.** The `smallestai` LiveKit plugin sends `eou_timeout_ms=100` unless you set it. Set `eou_timeout_ms=1000` for Hindi and English conversational audio.

**Error bodies on `POST /waves/v1/stt/`.** The API reference now shows the error bodies the server returns. A missing `Authorization` header returns `{ "message": "..." }`. Other authentication errors, and plan, credit, and rate-limit errors, return `{ "error": "..." }`. Validation and transcription errors return `{ "status": "error", "message": "..." }` with optional `errors`, `error_code`, `code`, `language`, and `region`.

→ [Finalization and Endpointing](/models/speech-to-text/features/endpointing)

## Pulse STT: European language codes removed from public batch enum

> The 14 European codes previously listed on the pre-recorded (batch) endpoint (uk, pl, cs, sk, lv, et, ro, fi, sv, bg, hu, da, lt, mt) are removed from the cus

The 14 European codes previously listed on the pre-recorded (batch) endpoint (`uk`, `pl`, `cs`, `sk`, `lv`, `et`, `ro`, `fi`, `sv`, `bg`, `hu`, `da`, `lt`, `mt`) are removed from the customer-facing batch enum on `POST /waves/v1/stt/` and `POST /waves/v1/pulse/get_text`.

Batch now advertises 12 single-language codes: `en`, `hi`, `de`, `es`, `ru`, `it`, `fr`, `nl`, `pt`, `zh`, `ja`, `ko`, plus the `multi-eu`, `multi-asian`, and `multi-indic` regional aggregators.

Streaming coverage is unchanged (21 single-language codes plus `north_indic`, `multi-asian`, `multi-south-indic`).

If your integration was routing one of the 14 removed codes to batch, migrate to `multi-eu`, which auto-detects across European languages. None of the 14 codes are available on streaming.

## Pulse STT: European language codes removed from public batch enum

> The 14 European codes previously listed on the pre-recorded (batch) endpoint (uk, pl, cs, sk, lv, et, ro, fi, sv, bg, hu, da, lt, mt) are removed from the cus

The 14 European codes previously listed on the pre-recorded (batch) endpoint (`uk`, `pl`, `cs`, `sk`, `lv`, `et`, `ro`, `fi`, `sv`, `bg`, `hu`, `da`, `lt`, `mt`) are removed from the customer-facing batch enum on `POST /waves/v1/stt/` and `POST /waves/v1/pulse/get_text`.

Batch now advertises 12 single-language codes: `en`, `hi`, `de`, `es`, `ru`, `it`, `fr`, `nl`, `pt`, `zh`, `ja`, `ko`, plus the `multi-eu`, `multi-asian`, and `multi-indic` regional aggregators.

Streaming coverage is unchanged (21 single-language codes plus `north_indic`, `multi-asian`, `multi-south-indic`).

If your integration was routing one of the 14 removed codes to batch, migrate to `multi-eu`, which auto-detects across European languages. None of the 14 codes are available on streaming.

## Pulse STT: keyword boosting on the pre-recorded endpoint

> Keyword boosting now works on the pre-recorded HTTP endpoint (POST /waves/v1/pulse/get_text and POST /waves/v1/stt/?model=pulse).

Keyword boosting now works on the pre-recorded HTTP endpoint (`POST /waves/v1/pulse/get_text` and `POST /waves/v1/stt/?model=pulse`). Pass `keywords` as a query parameter with the same `KEYWORD:INTENSIFIER` format as the realtime WebSocket. Max 100 keywords per request; sending more returns `400` with `keywords too large (max 100)` in `errors[]`.

**Intensifier guidance updated.** Default `1`, recommended `1` to `3`. Above `10` is not recommended: higher values increase the chance of hallucinating the keyword when it was not spoken.

**Client tip.** Reuse the same keyword list across requests when the vocabulary is stable.

→ [Keyword Boosting](/models/speech-to-text/features/keyword-boosting)

## Pulse STT: diarization response shape correction and rewrite

> Fixed two things the Speaker diarization page and the streaming WS spec got wrong, plus tightened the page for developers rather than marketing.

Fixed two things the [Speaker diarization](/models/speech-to-text/features/diarization) page and the streaming WS spec got wrong, plus tightened the page for developers rather than marketing.

**`speaker` is an integer on both batch and streaming.** The page previously claimed the batch API returned string labels (`speaker_0`, `speaker_1`, ...) while streaming returned integers. Live: both return integers, zero-indexed. Example values in the docs and the response-fields table are updated. The `stt-live-ws.yaml` spec had `speaker` typed as `string` on the streaming word/utterance schemas; corrected to `integer` with `example: 0` on both.

**`speaker_confidence` appears on batch too.** The table previously claimed the field was streaming-only. Live: batch responses also include `speaker_confidence` (0.0–1.0) on `words[]`. Corrected.

**Interim streaming events do not carry `words[]`.** New section clarifies that speaker attribution appears only on `is_final: true` transcription events. Live captions that rely on interim events should hold the speaker from the last final event and update on the next one.

**Pulse Pro carve-out.** `?model=pulse-pro&diarize=true` returns a transcript-only response with no `speaker` fields and no `utterances[]`. Now documented explicitly on the page.

Every claim on the page is verified against live production. Reproduce with `python3 scripts/spec-live-tests/diarization_live_test.py` (new); 16/16 checks pass.

## Improved Pulse STT release - stronger repeated-entity dictation and background-speaker handling

> Pulse STT has been updated with an improved English model.

Pulse STT has been updated with an improved English model. The Pulse Performance page (`/models/documentation/speech-to-text-pulse/benchmarks/performance`) has been refreshed with new WER for the standard public benchmarks - **FLEURS**, **HuggingFace ESB**, and **WildASR** - and improves on every one of them. WildASR sees the largest single-benchmark move, driven by big gains on the harder degradation subsets (Far-field, Clipping, Reverberation, Noise Gap).

Beyond the public benchmarks, two production-facing failure modes see substantial gains:

* **Repeated-entity dictation** - spoken digit sequences, alphanumeric IDs, PAN / credit-card / reference-number utterances, and any content with long runs of the same token. WER on this class of audio drops by **54% relatively** vs. the previous model. This is the same failure mode where earlier releases would collapse `"3 3 3 3 3 3 3"` down to a shorter run; the new model preserves run length far more consistently.

* **Background-speaker robustness** - audio where a secondary speaker or crosstalk is audible under the primary speaker. WER on this class also drops by **61% relatively**, so multi-speaker call recordings and open-office capture behave much closer to clean single-speaker audio than before.

These improvements are live behind the existing Pulse STT streaming endpoints - no API changes, no request-side flags to flip.

## Pulse STT streaming: external_session_id correlation parameter

> The streaming STT WebSocket now documents the external_session_id query parameter (available on both /waves/v1/stt/live and /waves/v1/pulse/get_text).

The streaming STT WebSocket now documents the `external_session_id` query parameter (available on both `/waves/v1/stt/live` and `/waves/v1/pulse/get_text`).

Pass your own session identifier — e.g. a call or conversation ID — and it is stored on the session's usage/analytics record, so you can join STT usage back to your own systems. Only alphanumeric characters, hyphens, underscores, and dots allowed; max 128 characters. It never affects transcription. Mirrors the TTS `session_id` correlation contract.

**Not echoed back.** Unlike the TTS WebSocket (where a client-provided `session_id` is returned as `external_session_id` on every response), STT responses carry only the server-generated `session_id`. Keep your own mapping client-side if you need per-message correlation.

→ [Pulse STT WebSocket reference](/api-reference/models/speech-to-text/speech-to-text)

## Pulse STT: European streaming Beta, redaction English/Hindi-only

**European streaming languages tagged Beta.** The 7 European codes on the streaming endpoint (`de`, `es`, `ru`, `it`, `fr`, `nl`, `pt`) now carry a Beta badge on the Pulse model card.

**Pre-recorded language list trimmed.** Removed `da`, `lv`, `et`, `mt` from the batch language table (not offered on batch). New counts: 22 pre-recorded, 31 total, 21 streaming unchanged. `multi-eu` aggregator now covers 17 European codes and is pre-recorded only.

**Redaction language scope.** `redact_pii` and `redact_pci` are supported on `language=en` and `language=hi` only. Set on other language codes, the API accepts the flag but does not reliably redact. Documented on the [Redaction](/models/speech-to-text/features/redaction) page and in the `stt-live-ws` spec.

## Pulse STT: European streaming Beta, redaction English/Hindi-only

**European streaming languages tagged Beta.** The 7 European codes on the streaming endpoint (`de`, `es`, `ru`, `it`, `fr`, `nl`, `pt`) now carry a Beta badge on the Pulse model card.

**Pre-recorded language list trimmed.** Removed `da`, `lv`, `et`, `mt` from the batch language table (not offered on batch). New counts: 22 pre-recorded, 31 total, 21 streaming unchanged. `multi-eu` aggregator now covers 17 European codes and is pre-recorded only.

**Redaction language scope.** `redact_pii` and `redact_pci` are supported on `language=en` and `language=hi` only. Set on other language codes, the API accepts the flag but does not reliably redact. Documented on the [Redaction](/models/speech-to-text/features/redaction) page and in the `stt-live-ws` spec.

## Pulse STT documentation cleanup, multi-indic restored, redact language scope

> Small STT docs corrections surfaced by customer reports and post-#251 audit.

Small STT docs corrections surfaced by customer reports and post-#251 audit.

**`multi-indic` restored on pre-recorded.** Documented as a pre-recorded aggregator (India region only) covering `en`, `hi`, `gu`, `mr`, `bn`, `or`. Streaming still uses `north_indic` for the same coverage; `multi-indic` is not supported on the streaming endpoint. Aligns the docs with what the platform actually accepts on `POST /waves/v1/stt/`.

**`redact_pii` / `redact_pci` language scope.** Param descriptions now note that redaction is currently effective only on `en` and `hi`. Setting these params `true` on other languages is accepted at the API layer but does not redact. Applies to both the streaming WebSocket (`stt-live-ws-overrides.yml`) and the pre-recorded REST endpoint (`stt-openapi.yaml`).

Also picks up small drift/miswritten-copy fixes across STT non-model-card pages: corrected Pulse TTFT number to \~150 ms (was 64 ms / "Sub-100 ms"), removed a stale `emotions` mention in the diarization enrichment list, and fixed a webhook sample payload that still showed `gender` + `emotions`.

→ [Pulse STT overview](/models/speech-to-text/overview)

## Pulse STT endpointing now on by default

> Pulse STT streaming finalizes each turn on trailing silence by default.

Pulse STT streaming finalizes each turn on trailing silence by default. When the speaker stops, the server emits `is_final: true` promptly instead of waiting for the model's internal finalization cadence. Transcript accuracy is unchanged.

To restore the previous finalization behavior, set `endpointing=false` on the WebSocket URL:

```
wss://api.smallest.ai/waves/v1/stt/live?model=pulse&endpointing=false
```

→ [Endpointing](/models/speech-to-text/features/endpointing)

## Pulse STT Beta tag on South Indian streaming languages

> Tamil, Telugu, Kannada, and Malayalam (plus the multi-south-indic aggregator) are now marked Beta on the Pulse model card and the STT overview.

Tamil, Telugu, Kannada, and Malayalam (plus the `multi-south-indic` aggregator) are now marked **Beta** on the [Pulse model card](/model-cards/speech-to-text/pulse) and the [STT overview](/models/speech-to-text/overview). Accuracy improvements are ongoing. These codes remain India-region only (`wss://api.smallest.ai/waves/v1/stt/live`); the US endpoint continues to reject them with `LANGUAGE_NOT_ENABLED_IN_REGION`.

Also cleaned up stale `multi` (bare) language references in the pulse-stt-ws-overrides.yml enum, the pulse-stt API reference, and the voice-cloning-openapi.yaml language example. `multi` was replaced by the current `multi-eu` / `multi-asian` / `north_indic` / `multi-south-indic` aggregators when the streaming enum was revised. Pre-recorded aggregators (`multi-eu`, `multi-asian`) are unchanged.

→ [Pulse model card](/model-cards/speech-to-text/pulse#supported-languages)

## Pulse STT Beta tag on South Indian streaming languages

> Tamil, Telugu, Kannada, and Malayalam (plus the multi-south-indic aggregator) are now marked Beta on the Pulse model card and the STT overview.

Tamil, Telugu, Kannada, and Malayalam (plus the `multi-south-indic` aggregator) are now marked **Beta** on the [Pulse model card](/model-cards/speech-to-text/pulse) and the [STT overview](/models/speech-to-text/overview). Accuracy improvements are ongoing. These codes remain India-region only (`wss://api.smallest.ai/waves/v1/stt/live`); the US endpoint continues to reject them with `LANGUAGE_NOT_ENABLED_IN_REGION`.

Also cleaned up stale `multi` (bare) language references in the pulse-stt-ws-overrides.yml enum, the pulse-stt API reference, and the voice-cloning-openapi.yaml language example. `multi` was replaced by the current `multi-eu` / `multi-asian` / `north_indic` / `multi-south-indic` aggregators when the streaming enum was revised. Pre-recorded aggregators (`multi-eu`, `multi-asian`) are unchanged.

→ [Pulse model card](/model-cards/speech-to-text/pulse#supported-languages)

## Pulse STT - streaming now supports ta / te / kn / ml + `multi-south-indic` aggregator (India region)

> The Pulse streaming Speech-to-Text API now supports four South Indian languages: Tamil (ta), Telugu (te), Kannada (kn), and Malayalam (ml).

The Pulse streaming Speech-to-Text API now supports four South Indian languages: Tamil (`ta`), Telugu (`te`), Kannada (`kn`), and Malayalam (`ml`). The `multi-south-indic` aggregator is also available for unknown South Indian audio - it auto-detects across the same four-language set plus English code-switching.

These five language values are served from the **India region only** - connect to `wss://api.smallest.ai/waves/v1/stt/live?model=pulse` (or the legacy `wss://api.smallest.ai/waves/v1/pulse/get_text`). Requesting any of them on the US host (`wss://api.us.smallest.ai/...`) returns an error:

```json
{
  "type": "error",
  "error_code": "LANGUAGE_NOT_ENABLED_IN_REGION",
  "message": "Language 'multi-south-indic' has not been enabled in this region. Please contact support to request access."
}
```

All other Pulse streaming parameters work as before: `word_timestamps`, `diarize`, `redact_pii`, `redact_pci`, `punctuate`, `capitalize`, `itn_normalize`, sample rates 8000/16000/22050/24000/44100/48000, encodings `linear16` (default) / `linear32` / `alaw` / `mulaw` / `opus` / `ogg_opus`.

**Pre-recorded (batch) is unchanged.** These four languages are not enabled on the batch endpoint - use the streaming endpoint.

→ [Pulse model card - Supported Languages](/model-cards/speech-to-text/pulse)
→ [Speech-to-Text overview](/models/speech-to-text/overview)

## Pulse STT - streaming now supports ja / yue / zh / ko + `multi-asian` aggregator (US region)

> The Pulse streaming Speech-to-Text API now supports four Asian languages: Japanese (ja), Cantonese (yue), Mandarin (zh), and Korean (ko).

The Pulse streaming Speech-to-Text API now supports four Asian languages: Japanese (`ja`), Cantonese (`yue`), Mandarin (`zh`), and Korean (`ko`). The `multi-asian` aggregator is also available for unknown East Asian audio - it auto-detects across the same four-language set.

These five language values are served from the US region only - connect to `wss://api.us.smallest.ai/waves/v1/stt/live?model=pulse` (or the legacy `wss://api.us.smallest.ai/waves/v1/pulse/get_text`) instead of the default `wss://api.smallest.ai/...` host. Requesting any of them on the default (ap-south-1) host closes the connection without a transcription frame.

All other Pulse streaming parameters work as before: `punctuate`, `capitalize`, `numerals`, `word_timestamps`, `diarize`, `redact_pii`, `redact_pci`, sample rates 8000/16000/22050/24000/44100/48000, encodings `linear16` (default) / `linear32` / `alaw` / `mulaw` / `opus` / `ogg_opus`.

**Pre-recorded (batch) is unchanged.** Cantonese (`yue`) is not enabled on the batch endpoint at all; Japanese, Mandarin, and Korean behave separately from streaming and are not formally supported on batch - use the streaming endpoint for these four languages.

Example response frame:

```json
{
  "type": "transcription",
  "transcript": "讀書要從薄到厚再從厚到薄",
  "is_final": true,
  "is_last": false,
  "language": "zh"
}
```

→ [Pulse model card - Supported Languages](/model-cards/speech-to-text/pulse)
→ [Speech-to-Text overview](/models/speech-to-text/overview)

## Streaming vs Pre-Recorded language framing across Pulse docs

> The Pulse STT supported-languages story is now consistent end-to-end.

The Pulse STT supported-languages story is now consistent end-to-end. Each surface labels its mode and points readers at the right neighbour:

* **Overview page** (`/models/documentation/speech-to-text-pulse/overview`): split the single Supported Languages table into two - Streaming (Real-Time, WebSocket) and Non-Streaming (Pre-Recorded, HTTP). Same 39-language set in both, matching the [Pulse model card](/model-cards/speech-to-text/pulse).
* **API Reference - Pre-Recorded (HTTP)**: `language` query param now declares an explicit `enum` of 39 single-language codes (`en`, `hi`, `es`, ...) plus `multi-eu` / `multi-indic` aggregators. Description labels the endpoint as **Pre-Recorded** and points at the streaming WebSocket endpoint for the live mode.
* **API Reference - Streaming (WebSocket)**: same enum and labelling, applied via the AsyncAPI override so it renders correctly. The page-top Note labels it as **Streaming** and points at the HTTP endpoint for pre-recorded.

No protocol or wire-level change - purely a docs/spec consolidation.

## Pulse STT WebSocket - `finalize` operation restored on the API reference

> The `sendFinalize` operation (and its `{"type":"finalize"}` payload) was silently missing from the rendered API reference page due to a spec-layer message-key mismatch. Restored alongside `sendCloseStream`. Wire behavior unchanged - both control messages have been accepted by the server since launch.

The Pulse STT WebSocket API reference now correctly renders **both** control-message operations on docs.smallest.ai:

* `sendFinalize` - payload `{"type":"finalize"}` - flushes the current audio buffer and emits an `is_final: true` transcript while keeping the session open. Useful for per-turn finalization in agentic pipelines.
* `sendCloseStream` - payload `{"type":"close_stream"}` - flushes any buffered audio, emits the terminal `is_final: true` + `is_last: true` transcript, then closes the session.

**No wire change** - the server has accepted both control messages since launch (`StreamControlType.FINALIZE` and `STREAM_CONTROL_TYPE_END` are sibling enum values in the internal gRPC schema, and the WS controller has explicit handlers for both `parsed.type === "finalize"` and `parsed.type === "close_stream"`). Only the rendered documentation was incomplete.

**Root cause** for the doc reviewers who want to know: the v4 docs override (`fern/apis/waves-v4/overrides/pulse-stt-ws-overrides.yml`) and the SDK override (`fern/apis/waves/asyncapi/pulse-stt-ws-overrides.yml`) both used unprefixed message keys (`audioData.message`, `finalizeSignal.message`, …) while the base spec used a `pulse*` prefix (`pulseAudioData.message`, `pulseFinalizeSignal.message`, …). Fern's merge silently dropped one of the three send operations when it couldn't resolve refs cleanly. Convention going forward: **override message keys must be identical to the base spec's keys** - every other Waves spec layer (TTS WS, Lightning v3.1 WS) already follows this rule.

**Migration:** nothing for customers. The server-side contract has always allowed both control messages; this is purely a docs render fix.

**Also clarified:** the ITN feature page's "Recommended Setup for Agentic Use Cases" section had been recommending `close_stream` per utterance, which is correct for single-shot transcription but misleading for the multi-turn voice agents the section is named after. Split the recommendation into two paths: `finalize` per user turn (session stays open, lowest latency between turns) vs `close_stream` at end of session (terminal). The Python example was split into two snippets that match these two patterns directly.

**Also clarified - clearer signal framing:** the `finalize` and `close_stream` control messages are now documented as a *turn-boundary signal* vs *session-end signal* on the API ref (operation summaries) and in the ITN feature page. The base-spec operation summaries used to read "Flush current audio buffer" / "End the audio stream" - accurate but not actionable. They now say what the signal *does to the session lifecycle*: keep listening vs hang up the socket.

**CI lock-in:** `spec_drift_check.py` was extended with an AsyncAPI override-key parity check. Any new override whose `channels.<chan>.messages.<KEY>` or `operations.<KEY>` doesn't exist in the base will fail the gate. Existing deprecated-spec drift (Lightning v2, the legacy `/streaming-tts/stream` route) is allow-listed with a documented rationale and tracked separately. A new post-deploy smoke check (`docs_render_smoke.py`, wired into `publish-docs.yml`) re-fetches the rendered docs after every push to main and asserts every expected operation appears - the same check would have caught this exact bug the day it deployed.

## Pulse STT WebSocket - `finalize` operation restored on the API reference

> The `sendFinalize` operation (and its `{"type":"finalize"}` payload) was silently missing from the rendered API reference page due to a spec-layer message-key mismatch. Restored alongside `sendCloseStream`. Wire behavior unchanged - both control messages have been accepted by the server since launch.

The Pulse STT WebSocket API reference now correctly renders **both** control-message operations on docs.smallest.ai:

* `sendFinalize` - payload `{"type":"finalize"}` - flushes the current audio buffer and emits an `is_final: true` transcript while keeping the session open. Useful for per-turn finalization in agentic pipelines.
* `sendCloseStream` - payload `{"type":"close_stream"}` - flushes any buffered audio, emits the terminal `is_final: true` + `is_last: true` transcript, then closes the session.

**No wire change** - the server has accepted both control messages since launch (`StreamControlType.FINALIZE` and `STREAM_CONTROL_TYPE_END` are sibling enum values in the internal gRPC schema, and the WS controller has explicit handlers for both `parsed.type === "finalize"` and `parsed.type === "close_stream"`). Only the rendered documentation was incomplete.

**Root cause** for the doc reviewers who want to know: the v4 docs override (`fern/apis/waves-v4/overrides/pulse-stt-ws-overrides.yml`) and the SDK override (`fern/apis/waves/asyncapi/pulse-stt-ws-overrides.yml`) both used unprefixed message keys (`audioData.message`, `finalizeSignal.message`, …) while the base spec used a `pulse*` prefix (`pulseAudioData.message`, `pulseFinalizeSignal.message`, …). Fern's merge silently dropped one of the three send operations when it couldn't resolve refs cleanly. Convention going forward: **override message keys must be identical to the base spec's keys** - every other Waves spec layer (TTS WS, Lightning v3.1 WS) already follows this rule.

**Migration:** nothing for customers. The server-side contract has always allowed both control messages; this is purely a docs render fix.

**Also clarified:** the ITN feature page's "Recommended Setup for Agentic Use Cases" section had been recommending `close_stream` per utterance, which is correct for single-shot transcription but misleading for the multi-turn voice agents the section is named after. Split the recommendation into two paths: `finalize` per user turn (session stays open, lowest latency between turns) vs `close_stream` at end of session (terminal). The Python example was split into two snippets that match these two patterns directly.

**Also clarified - clearer signal framing:** the `finalize` and `close_stream` control messages are now documented as a *turn-boundary signal* vs *session-end signal* on the API ref (operation summaries) and in the ITN feature page. The base-spec operation summaries used to read "Flush current audio buffer" / "End the audio stream" - accurate but not actionable. They now say what the signal *does to the session lifecycle*: keep listening vs hang up the socket.

**CI lock-in:** `spec_drift_check.py` was extended with an AsyncAPI override-key parity check. Any new override whose `channels.<chan>.messages.<KEY>` or `operations.<KEY>` doesn't exist in the base will fail the gate. Existing deprecated-spec drift (Lightning v2, the legacy `/streaming-tts/stream` route) is allow-listed with a documented rationale and tracked separately. A new post-deploy smoke check (`docs_render_smoke.py`, wired into `publish-docs.yml`) re-fetches the rendered docs after every push to main and asserts every expected operation appears - the same check would have caught this exact bug the day it deployed.

## Unified Speech-to-Text endpoint, Pulse Pro model

> The Speech-to-Text API now lives at the unified path /waves/v1/stt/, mirroring the unified TTS shape.

The Speech-to-Text API now lives at the unified path `/waves/v1/stt/`, mirroring the unified TTS shape. The model is selected via the `?model=` query parameter. Two models are live today:

* `?model=pulse`: multilingual (17 streaming + 26 pre-recorded languages), HTTP + WebSocket streaming.
* `?model=pulse-pro`: leaderboard-ranked English STT (5.42% ESB avg WER, tied #2 on the public Open ASR Leaderboard). HTTP only. Pass `webhook_url` for long files.

**Customer pricing (Standard plan):**

* Pulse, streaming (WebSocket): \$0.006 / minute
* Pulse, non-streaming (HTTP): \$0.0035 / minute
* Pulse Pro, non-streaming (HTTP): \$0.004 / minute

Standard plan rate limits default to 25 RPM per model and 100 concurrent WebSocket sessions. Enterprise is unlimited and configurable per-customer.

The existing endpoints (`POST /waves/v1/pulse/get_text` and `WS /waves/v1/pulse/get_text`) continue to work alongside the new unified path. New integrations are encouraged to use `/waves/v1/stt/` since it carries both models behind one path.

* [Pulse Pro model card](/model-cards/speech-to-text/pulse-pro)
* [Speech-to-Text quickstart](/models/speech-to-text/pre-recorded/quickstart) covers both models.

## Pulse STT - full feature parity over HTTP and WebSocket

> All Pulse STT features (speaker diarization, word timestamps, emotion, redaction, etc.) are now consistently available on both HTTP (pre-recorded) and WebSocket (realtime) modes.

The full Pulse STT feature set is now available on **both** HTTP (pre-recorded) and WebSocket (realtime) modes with consistent flag names and response shapes.

**What this means in practice:**

* Same query parameters and request body fields work the same way on `POST /waves/v1/pulse/get_text` (pre-recorded) and `wss://api.smallest.ai/waves/v1/pulse/get_text` (realtime).
* Speaker diarization, word timestamps, sentence-level utterances, emotion detection, gender detection, keyword boosting, redaction, punctuation, and inverse text normalisation all behave identically across modes.
* The full per-feature documentation is on the [Features](/models/speech-to-text/features/word-timestamps) pages.

**Migration:** no action - additive. Existing integrations on either mode continue working.

_Showing the 20 most recent of 32 entries. Append `/llms.txt` to the changelog URL for the complete index._