Skip to navigation

End-of-Utterance Timeout

Control how long Pulse waits after speech ends before finalizing the transcript.
View as Markdown
Real-Time

End-of-utterance (EOU) timeout controls how long the model waits in silence after a speaker stops talking before it flushes the transcript as final. Tuning this value lets you balance responsiveness against cutting users off mid-thought.

With endpointing=true (the default), eou_timeout_ms sets the trailing-silence window (600 ms if unset). With endpointing=false, it sets the model’s own end-of-utterance timer (800 ms if unset). See Finalization and Endpointing for every finalization trigger and the recommended parameters per orchestration.

How It Works

When speech pauses, Pulse starts a silence timer. If no additional speech is detected within the eou_timeout_ms window, the current transcript segment is returned with is_final: true.

  • Lower values: faster turn detection, but more likely to split natural pauses
  • Higher values: more tolerant of pauses, but slower finalization

Enabling EOU Timeout

Add eou_timeout_ms to your WebSocket connection query parameters. The value must be an integer from 100 to 10000. If unset, the window is 600 ms with endpointing=true (the default) and 800 ms with endpointing=false.

const url = new URL("wss://api.smallest.ai/waves/v1/stt/live?model=pulse");
url.searchParams.append("language", "en");
url.searchParams.append("encoding", "linear16");
url.searchParams.append("sample_rate", "16000");
url.searchParams.append("eou_timeout_ms", "300"); // fast turn-taking
const ws = new WebSocket(url.toString(), {
headers: {
Authorization: `Bearer ${API_KEY}`,
},
});

How to Tune It

For conversational voice agents, start at 1000 ms. Mid-sentence pauses in phone conversations often last around 800 ms, so shorter windows close the final before the speaker has finished.

  • Decrease only for short, one-word answers where speed matters more than sentence completeness
  • Increase for meeting or dictation workflows where speakers pause longer

Tuning Guide

ValueBehaviorBest for
300-600msFast. Splits sentences at natural mid-sentence pausesShort commands and one-word answers
800-1200msHolds a sentence together through natural pausesVoice agents, conversational AI
1500-2000msPatient. Waits through longer pausesMeeting transcription, dictation, accessibility
3000ms+Very patient - rarely flushes earlyLecture capture, users who pause frequently

Trade-offs

DimensionLow timeout (e.g. 300ms)Recommended starting point (1000ms)
Response speedFast. The final arrives soon after the speaker stopsAbout 1 s after the speaker stops
Turn accuracySplits sentences at natural mid-sentence pausesHolds a sentence together through pauses of up to about 1 s
Best forShort commands and one-word answersVoice agents and conversational AI

Example

A voice agent needs one final per caller sentence:

wss://api.smallest.ai/waves/v1/stt/live?model=pulse&language=en&encoding=linear16&sample_rate=16000&eou_timeout_ms=1000

A meeting transcription system should wait for natural pauses:

wss://api.smallest.ai/waves/v1/stt/live?model=pulse&language=en&encoding=linear16&sample_rate=16000&eou_timeout_ms=1500