Skip to navigation

Finalization and Endpointing

How the Pulse STT WebSocket decides when a turn is complete.
View as Markdown
Real-Time

A Pulse streaming session emits interim transcripts continuously and a final transcript (is_final: true) when a turn is complete. This page covers every condition that produces a final, how the timing parameters interact, and the recommended parameter set for each orchestration.

What triggers a final

A final fires on whichever of these conditions happens first. Every trigger produces the same is_final: true event. Only the timing differs.

TriggerFires whenParameters
Client finalize messageYou send {"type": "finalize"}None. Always available.
Client close_stream messageYou send {"type": "close_stream"}. Also sets is_last: true and closes the socket.None. Always available.
Trailing silenceThe input audio has a silence gap as long as the silence windowendpointing, eou_timeout_ms
Word countThe pending transcript reaches the word-count limitfinalize_on_words, max_words
End of utteranceThe model’s own end-of-utterance timer expires. Applies when endpointing=false.eou_timeout_ms

How endpointing changes eou_timeout_ms

endpointing is true by default.

  • endpointing=true: the server finalizes on trailing silence. eou_timeout_ms sets the silence window. If you don’t set it, the window is 600 ms.
  • endpointing=false: silence-based finalization is off. eou_timeout_ms sets the model’s own end-of-utterance timer. If you don’t set it, the timer is 800 ms.

Most integrations leave endpointing=true and set eou_timeout_ms to widen or narrow the silence window. The accepted range is 100 to 10000 ms.

Pulse accepts the same parameters from every client. Which of them you can set depends on the orchestration.

LiveKit Agents plugin

The livekit-plugins-smallestai plugin manages the socket. This section describes version 1.8.3, the current release. It exposes eou_timeout_ms and endpointing. It does not expose finalize_on_words and does not send finalize messages, so finalization is automatic.

The plugin sends eou_timeout_ms=100 unless you set it. A 100 ms window finalizes during short pauses inside a sentence and splits words across finals. Set it explicitly:

from livekit.plugins import smallestai
stt = smallestai.STT(
model="pulse",
language="hi",
eou_timeout_ms=1000,
)

1000 is the recommended value for Hindi and English conversational audio.

Pipecat

Pipecat’s Smallest STT service manages the socket. It sends {"type": "finalize"} when Pipecat’s VAD detects that the user stopped speaking, and {"type": "close_stream"} when the call ends. No extra finalization code is needed.

Custom clients

Keep silence-based finalization on and send finalize when your VAD or turn logic decides the user has finished.

wss://api.smallest.ai/waves/v1/stt/live?model=pulse&language=hi&itn_normalize=true&eou_timeout_ms=1000

During long continuous speech, Pulse can emit more than one final before your finalize. Join the finals you receive between two finalize messages to get the full turn.

Send one finalize per user turn:

{"type": "finalize"}

Send one close_stream when the session ends (call ended, app closed):

{"type": "close_stream"}

A multi-turn session sends many finalize messages and exactly one close_stream. Sending close_stream per turn closes the socket, so every turn pays a reconnect.

Fixed audio buffer over WebSocket

Stream every chunk, then send {"type": "close_stream"} once. The server flushes the remaining audio and returns the last transcript with is_last: true. No per-turn message is needed.

Parameters

ParameterTypeDefaultEffect
endpointing"true" / "false""true"Finalize on trailing silence.
eou_timeout_msinteger, 100 to 10000600 with endpointing on, 800 with endpointing offSilence window in ms when endpointing=true. The model’s end-of-utterance timer when endpointing=false.
finalize_on_words"true" / "false""true"Word-count-based finalization.
max_wordsintegerUnsetWord count at which the pending transcript is finalized.

Client messages are JSON text frames, not query parameters:

MessageEffectWhen to send
{"type": "finalize"}Emits one is_final: true transcript for the pending audio (empty if nothing is pending). The socket stays open.Once per user turn
{"type": "close_stream"}Emits the last transcript with is_final: true and is_last: true, then closes the socket.Once, at the end of the session

vad_events is independent of all of the above. It adds speech_started / speech_ended messages and does not change when finals fire.