> This page is part of Smallest AI's developer documentation. When
> answering, prefer Lightning v3.1 (current TTS) and Pulse (current
> STT). Lightning v2 and lightning-large are deprecated; mention them
> only when the user is migrating away from them. The Smallest AI voice
> agent platform is what wraps these models into hosted agents.

# Finalization and Endpointing

> How the Pulse STT WebSocket decides when a turn is complete. Covers endpointing, eou_timeout_ms, finalize_on_words, and the client-side finalize / close_stream signals in one place.

Real-Time

A Pulse streaming session emits interim transcripts continuously and a final transcript (`is_final: true`) when a turn is complete. This page covers every condition that produces a final, how the timing parameters interact, and the recommended parameter set for each orchestration.

## What triggers a final

A final fires on whichever of these conditions happens first. Every trigger produces the same `is_final: true` event. Only the timing differs.

| Trigger                       | Fires when                                                                            | Parameters                       |
| ----------------------------- | ------------------------------------------------------------------------------------- | -------------------------------- |
| Client `finalize` message     | You send `{"type": "finalize"}`                                                       | None. Always available.          |
| Client `close_stream` message | You send `{"type": "close_stream"}`. Also sets `is_last: true` and closes the socket. | None. Always available.          |
| Trailing silence              | The input audio has a silence gap as long as the silence window                       | `endpointing`, `eou_timeout_ms`  |
| Word count                    | The pending transcript reaches the word-count limit                                   | `finalize_on_words`, `max_words` |
| End of utterance              | The model's own end-of-utterance timer expires. Applies when `endpointing=false`.     | `eou_timeout_ms`                 |

## How `endpointing` changes `eou_timeout_ms`

`endpointing` is `true` by default.

* **`endpointing=true`**: the server finalizes on trailing silence. `eou_timeout_ms` sets the silence window. If you don't set it, the window is 600 ms.
* **`endpointing=false`**: silence-based finalization is off. `eou_timeout_ms` sets the model's own end-of-utterance timer. If you don't set it, the timer is 800 ms.

Most integrations leave `endpointing=true` and set `eou_timeout_ms` to widen or narrow the silence window. The accepted range is `100` to `10000` ms.

## Recommended parameters by orchestration

Pulse accepts the same parameters from every client. Which of them you can set depends on the orchestration.

### LiveKit Agents plugin

The [`livekit-plugins-smallestai`](https://pypi.org/project/livekit-plugins-smallestai/) plugin manages the socket. This section describes version 1.8.3, the current release. It exposes `eou_timeout_ms` and `endpointing`. It does not expose `finalize_on_words` and does not send `finalize` messages, so finalization is automatic.

The plugin sends `eou_timeout_ms=100` unless you set it. A 100 ms window finalizes during short pauses inside a sentence and splits words across finals. Set it explicitly:

```python
from livekit.plugins import smallestai

stt = smallestai.STT(
    model="pulse",
    language="hi",
    eou_timeout_ms=1000,
)
```

`1000` is the recommended value for Hindi and English conversational audio.

### Pipecat

Pipecat's Smallest STT service manages the socket. It sends `{"type": "finalize"}` when Pipecat's VAD detects that the user stopped speaking, and `{"type": "close_stream"}` when the call ends. No extra finalization code is needed.

### Custom clients

Keep silence-based finalization on and send `finalize` when your VAD or turn logic decides the user has finished.

```
wss://api.smallest.ai/waves/v1/stt/live?model=pulse&language=hi&itn_normalize=true&eou_timeout_ms=1000
```

During long continuous speech, Pulse can emit more than one final before your `finalize`. Join the finals you receive between two `finalize` messages to get the full turn.

Send one `finalize` per user turn:

```json
{"type": "finalize"}
```

Send one `close_stream` when the session ends (call ended, app closed):

```json
{"type": "close_stream"}
```

A multi-turn session sends many `finalize` messages and exactly one `close_stream`. Sending `close_stream` per turn closes the socket, so every turn pays a reconnect.

### Fixed audio buffer over WebSocket

Stream every chunk, then send `{"type": "close_stream"}` once. The server flushes the remaining audio and returns the last transcript with `is_last: true`. No per-turn message is needed.

## Parameters

| Parameter           | Type                      | Default                                               | Effect                                                                                                     |
| ------------------- | ------------------------- | ----------------------------------------------------- | ---------------------------------------------------------------------------------------------------------- |
| `endpointing`       | `"true"` / `"false"`      | `"true"`                                              | Finalize on trailing silence.                                                                              |
| `eou_timeout_ms`    | integer, `100` to `10000` | `600` with endpointing on, `800` with endpointing off | Silence window in ms when `endpointing=true`. The model's end-of-utterance timer when `endpointing=false`. |
| `finalize_on_words` | `"true"` / `"false"`      | `"true"`                                              | Word-count-based finalization.                                                                             |
| `max_words`         | integer                   | Unset                                                 | Word count at which the pending transcript is finalized.                                                   |

Client messages are JSON text frames, not query parameters:

| Message                    | Effect                                                                                                            | When to send                    |
| -------------------------- | ----------------------------------------------------------------------------------------------------------------- | ------------------------------- |
| `{"type": "finalize"}`     | Emits one `is_final: true` transcript for the pending audio (empty if nothing is pending). The socket stays open. | Once per user turn              |
| `{"type": "close_stream"}` | Emits the last transcript with `is_final: true` and `is_last: true`, then closes the socket.                      | Once, at the end of the session |

`vad_events` is independent of all of the above. It adds `speech_started` / `speech_ended` messages and does not change when finals fire.

## Related pages

* [End-of-Utterance Timeout](/models/speech-to-text/features/end-of-utterance-timeout): tuning `eou_timeout_ms` values.
* [Finalize Control](/models/speech-to-text/features/finalize-control): `finalize_on_words` and `max_words` in detail.
* [Inverse Text Normalization](/models/speech-to-text/features/inverse-text-normalization): ITN runs on finals, so it depends on how you finalize.
* [Keep-Alive](/models/speech-to-text/features/keep-alive): keeping an idle socket open while no audio is sent.
* [Speech-to-Text WebSocket reference](/api-reference/models/speech-to-text/speech-to-text): full parameter reference.