Finalization and Endpointing
A Pulse streaming session emits interim transcripts continuously and a final transcript (is_final: true) when a turn is complete. This page covers every condition that produces a final, how the timing parameters interact, and the recommended parameter set for each orchestration.
What triggers a final
A final fires on whichever of these conditions happens first. Every trigger produces the same is_final: true event. Only the timing differs.
How endpointing changes eou_timeout_ms
endpointing is true by default.
endpointing=true: the server finalizes on trailing silence.eou_timeout_mssets the silence window. If you don’t set it, the window is 600 ms.endpointing=false: silence-based finalization is off.eou_timeout_mssets the model’s own end-of-utterance timer. If you don’t set it, the timer is 800 ms.
Most integrations leave endpointing=true and set eou_timeout_ms to widen or narrow the silence window. The accepted range is 100 to 10000 ms.
Recommended parameters by orchestration
Pulse accepts the same parameters from every client. Which of them you can set depends on the orchestration.
LiveKit Agents plugin
The livekit-plugins-smallestai plugin manages the socket. This section describes version 1.8.3, the current release. It exposes eou_timeout_ms and endpointing. It does not expose finalize_on_words and does not send finalize messages, so finalization is automatic.
The plugin sends eou_timeout_ms=100 unless you set it. A 100 ms window finalizes during short pauses inside a sentence and splits words across finals. Set it explicitly:
1000 is the recommended value for Hindi and English conversational audio.
Pipecat
Pipecat’s Smallest STT service manages the socket. It sends {"type": "finalize"} when Pipecat’s VAD detects that the user stopped speaking, and {"type": "close_stream"} when the call ends. No extra finalization code is needed.
Custom clients
Keep silence-based finalization on and send finalize when your VAD or turn logic decides the user has finished.
During long continuous speech, Pulse can emit more than one final before your finalize. Join the finals you receive between two finalize messages to get the full turn.
Send one finalize per user turn:
Send one close_stream when the session ends (call ended, app closed):
A multi-turn session sends many finalize messages and exactly one close_stream. Sending close_stream per turn closes the socket, so every turn pays a reconnect.
Fixed audio buffer over WebSocket
Stream every chunk, then send {"type": "close_stream"} once. The server flushes the remaining audio and returns the last transcript with is_last: true. No per-turn message is needed.
Parameters
Client messages are JSON text frames, not query parameters:
vad_events is independent of all of the above. It adds speech_started / speech_ended messages and does not change when finals fire.
Related pages
- End-of-Utterance Timeout: tuning
eou_timeout_msvalues. - Finalize Control:
finalize_on_wordsandmax_wordsin detail. - Inverse Text Normalization: ITN runs on finals, so it depends on how you finalize.
- Keep-Alive: keeping an idle socket open while no audio is sent.
- Speech-to-Text WebSocket reference: full parameter reference.