Skip to navigation

Pulse 2.0 launch: built-in end-of-turn detection, emotion, and gender

New model: Pulse 2.0. English-only, streaming-only speech to text on the unified live endpoint. Select it with ?model=pulse-2 on wss://api.smallest.ai/waves/v1/stt/live - the same endpoint Pulse streams on.

Built-in end-of-turn (EOT) detection. Finals close on turn boundaries instead of a fixed silence timeout or word count, so finals are fewer and longer. There are no tuning knobs. eou_timeout_ms, finalize_on_words, and max_words are ignored by default on model=pulse-2 - don’t send them. If you do, they still take effect and compete with end-of-turn detection. Finals closed by EOT carry eot_finalized: true.

Per-utterance emotion and gender, on by default. emotion_detection adds a six-class emotion object (neutral, happy, angry, sad, fear, disgust) with confidence and the full score distribution to every final. gender_detection adds a gender object (male, female) with confidence and scores. Both default to true; set either to false to opt out.

English only, streaming only. Any language other than en returns LANGUAGE_NOT_SUPPORTED_BY_MODEL and closes the socket. There is no pre-recorded/batch mode - for multilingual or batch transcription, use standard Pulse or Pulse Pro.

Pricing: $0.005 per minute of audio (Standard plan). Standard plan rate limits default to 100 concurrent WebSocket sessions, same as Pulse streaming.

→ Pulse 2.0 model card