Pulse 2.0
English-only streaming speech to text with built-in end-of-turn detection, per-utterance emotion, and per-utterance gender - all from one WebSocket. Start with the realtime quickstart.
Who it’s for
- Voice agents and live call analytics that need to know when the caller finished speaking and how they sounded, on every final.
- English-only streaming audio where built-in turn detection replaces a hand-tuned silence timeout.
- Not the pick for multilingual or batch audio. Use Pulse or Pulse Pro.
Model Overview
Key Capabilities
A built-in model decides when a speaker has finished their turn. Finals close on turn boundaries instead of a fixed ~10-word or 800 ms silence window, so finals are fewer and longer. No tuning knobs to configure.
Six-class emotion (neutral, happy, angry, sad, fear, disgust) with confidence and full score distribution on every final. On by default - opt out with emotion_detection=false.
Acoustic gender estimate (male, female) with confidence and scores on every final. On by default - opt out with gender_detection=false.
Speaker labels on words plus diarization segments, on by default.
Built-in redaction of personal data and payment-card information.
Per-word timing and custom-vocabulary boosting, same as the rest of the Pulse family.
Performance & Benchmarks
Best for: voice agents and live call analytics, where you need to know when the caller finished speaking and how they sounded.
Transcription accuracy (WER)
Unchanged from Pulse. The emotion and gender heads are a LoRA adapter plus classification heads added on top of the shared encoder, not a retrain of it, so recognition accuracy is identical to shipped Pulse - see the Pulse model card for WER benchmarks.
End-of-turn detection
Evaluated on livekit/eot-bench, English validation split (400 turns, 705 hold spans, 400 EOT spans). Every model streamed through its own live server and was scored against the same reference spans.
Sorted by false cutoffs at 300 ms, matching LiveKit’s own comparison table. Lower is better on every column; – means the model can’t reach that operating point.
Pulse EoT v1 ranks 3rd of 11 on false cutoffs at 600 ms and 3rd/4th on latency at the 5%/10% budgets - second only to LiveKit v1 among self-hosted models. Against Deepgram Flux specifically: it wins on false cutoffs at 600 ms (8.8% vs 9.9%) and latency at 5% (784 ms vs 1151 ms), and loses on false cutoffs at 300 ms and latency at 10%.
Signal-quality diagnostic (AUC / AP over the full span, independent of any operating-point choice):
2nd of 10 on raw signal quality. The gap between 2nd on discrimination and mid-pack on false cutoffs at the 300 ms budget is a timing effect, not a signal-quality one - the model knows the turn has ended, it just says so a little late.
Detect rate (correctly identifying a true turn-end at all, independent of timing) is best-in-class: 95.8% at the 5% false-cutoff operating point, vs LiveKit v1’s 91.0%.
Emotion detection
Evaluated on two corpora held out from training:
UA = unweighted accuracy, WA = weighted accuracy. Evaluated across nine emotion classes (anger, happiness, sadness, neutral, excitement, frustration, fear, surprise, disgust); the streaming API surfaces six classes per final - see Response Format.
Streaming introduces essentially no accuracy cost versus the direct (offline) model - per-corpus deltas average within about ±0.3 points, consistent with run-to-run noise rather than degradation.
Gender detection
95% F1, evaluated both intra-corpus and on held-out test sets.
Supported Languages
Pulse 2.0 is English-only.
Requesting any other language value returns a LANGUAGE_NOT_SUPPORTED_BY_MODEL error and closes the socket (see Error Frames). For multilingual or pre-recorded transcription, use standard Pulse (21 streaming + 12 pre-recorded languages).
Authentication
Include your API key in the Authorization header when opening the WebSocket:
Browsers can’t set headers on a WebSocket - mint a short-lived access token on your server and pass it as the api_key query parameter instead.
Request Parameters
Passed as query parameters on the WebSocket URL, for example:
emotion_detection / gender_detection are the exact param names - emotion=false or gender=false are silently ignored and both fields keep coming back with their defaults.
On pulse-2, finalization is driven by built-in end-of-turn detection, which has no tuning knobs. Pulse’s finalization params eou_timeout_ms, finalize_on_words, and max_words are ignored by default - don’t send them. If you do, they still take effect and compete with end-of-turn detection: for example, max_words=5 splits turns into short fragments and EOT fires less often. If you’re migrating from pulse, remove them from your client. See Finalization and Endpointing for how these work on Pulse.
An unknown model is rejected at the handshake with HTTP 400 INVALID_MODEL.
Control Messages
Sent from the client as JSON text frames on the same WebSocket:
Response Format
Every server message is a JSON frame. Interim (is_final: false) messages carry a running transcript string only; finals carry the full set of fields below, subject to which params you set.
Real final from prod (words trimmed to two):
Three things to keep in mind when integrating:
- Treat
emotionandgenderas optional. They are left out (notnull) when an utterance is too short or classification times out. - Expect fewer, longer finals. On a 25 s speech clip, Pulse 2.0 returned 5 finals, 4 of them
eot_finalized, each ending on a sentence or clause. - The last final has no
eot_finalized. It comes from closing the stream, marked byis_last: trueinstead.
Error Frames
Error frames always carry type: "error", error_code, and a human-readable message.
API Reference
Endpoint
See Transcribe (Realtime / WebSocket) for the full generated request/response schema, including this page’s params and fields.
There is no batch / pre-recorded Pulse 2.0. For high-volume English batch transcription, use Pulse Pro.
Throughput, Latency & Pricing
Pulse 2.0-specific latency numbers are being finalized. Pulse 2.0 runs on the same streaming infrastructure as Pulse - see Pulse’s Throughput, Latency & Pricing as a reference point in the meantime.
Customer pricing: $0.005 per minute of audio (Standard plan, streaming). Standard plan rate-limit default: 100 concurrent requests. Enterprise tier is unlimited and configurable per-customer.
Use Cases
Limitations
- English only. Use Pulse for every other language.
- Emotion and gender are not guaranteed on every final. They’re estimated from the voice of each utterance and can be left out on a very short utterance (under ~0.4 s) or a slow classification pass.
- Gender is a binary acoustic estimate of the voice, not an identity attribute. It describes how the speaker’s voice sounds to the model, not how the speaker identifies.
Safety & Compliance
Pulse 2.0 must not be used for:
- Recording or transcribing individuals without their explicit consent
- Surveillance, stalking, or any form of unauthorized monitoring
- Any illegal or unethical purposes
Additionally:
- Usage is monitored for policy compliance
- Gender output is an acoustic estimate, not an identity signal - do not use it to infer or record a speaker’s gender identity
- For compliance documentation (GDPR, SOC2, HIPAA), contact support@smallest.ai
FAQ
What is the difference between Pulse and Pulse 2.0?
Pulse 2.0 is English-only and streaming-only. On top of everything Pulse’s streaming mode supports, it adds built-in end-of-turn detection and per-utterance emotion and gender, on by default. For multilingual or batch transcription, use standard Pulse.
Why are finals fewer and longer on Pulse 2.0?
Pulse 2.0’s built-in end-of-turn model closes a final when it detects the speaker has finished their turn, rather than on a fixed ~10-word or 800 ms silence window. On a 25 s speech clip this produced 5 finals (4 of them eot_finalized).
Do I need to change my client code to turn off emotion or gender?
Add emotion_detection=false or gender_detection=false to the WebSocket URL - both default to true. Double-check the param name: emotion=false or gender=false are silently ignored and the fields keep coming back.
Why is emotion or gender missing from a final?
Why is emotion or gender missing from a final?
Both fields are best-effort per-utterance classifications. They’re omitted (not sent as null) when the utterance is too short (under ~0.4 s) for the classifier, or if classification doesn’t complete in time. Treat both fields as optional in your parsing.
What does eot_finalized mean, and why is it missing on the last final?
eot_finalized: true means the built-in end-of-turn model decided the speaker had finished talking. The very last final of a session comes from the client closing the stream (close_stream) or calling finalize, not from end-of-turn detection - that final is marked with is_last: true instead and has no eot_finalized field.
Can I still tune endpointing, eou_timeout_ms, or finalize_on_words on Pulse 2.0?
You shouldn’t. Pulse 2.0 replaces silence- and word-count-based finalization with built-in end-of-turn detection, which has no tuning knobs. eou_timeout_ms, finalize_on_words, and max_words are ignored by default - don’t send them. If you do, they still take effect and compete with end-of-turn detection, so you’ll get shorter, fragmented finals.