Skip to navigation

Pulse 2.0

Model card for Pulse 2.0, English-only streaming speech to text with built-in turn detection.
View as Markdown
Latest Release

English-only streaming speech to text with built-in end-of-turn detection, per-utterance emotion, and per-utterance gender - all from one WebSocket. Start with the realtime quickstart.

Who it’s for

  • Voice agents and live call analytics that need to know when the caller finished speaking and how they sounded, on every final.
  • English-only streaming audio where built-in turn detection replaces a hand-tuned silence timeout.
  • Not the pick for multilingual or batch audio. Use Pulse or Pulse Pro.

Model Overview

Developed bySmallest AI
Model typeSpeech-to-Text Streaming
Model IDpulse-2
LanguagesEnglish (en) only
Audio input formatslinear16, linear32, alaw, mulaw, opus, ogg_opus
Pricing
(Standard Plan)
$0.005/min
Concurrency
(Standard Plan)
100 concurrent requests
Recommended Sample Rate16,000 Hz
Endpointwss://api.smallest.ai/waves/v1/stt/live?model=pulse-2

Key Capabilities

End-of-Turn Detection

A built-in model decides when a speaker has finished their turn. Finals close on turn boundaries instead of a fixed ~10-word or 800 ms silence window, so finals are fewer and longer. No tuning knobs to configure.

Per-Utterance Emotion

Six-class emotion (neutral, happy, angry, sad, fear, disgust) with confidence and full score distribution on every final. On by default - opt out with emotion_detection=false.

Per-Utterance Gender

Acoustic gender estimate (male, female) with confidence and scores on every final. On by default - opt out with gender_detection=false.

Speaker Diarization

Speaker labels on words plus diarization segments, on by default.

PII / PCI Redaction

Built-in redaction of personal data and payment-card information.

Word Timestamps & Keyword Boosting

Per-word timing and custom-vocabulary boosting, same as the rest of the Pulse family.


Performance & Benchmarks

Best for: voice agents and live call analytics, where you need to know when the caller finished speaking and how they sounded.

Transcription accuracy (WER)

Unchanged from Pulse. The emotion and gender heads are a LoRA adapter plus classification heads added on top of the shared encoder, not a retrain of it, so recognition accuracy is identical to shipped Pulse - see the Pulse model card for WER benchmarks.

End-of-turn detection

Evaluated on livekit/eot-bench, English validation split (400 turns, 705 hold spans, 400 EOT spans). Every model streamed through its own live server and was scored against the same reference spans.

Sorted by false cutoffs at 300 ms, matching LiveKit’s own comparison table. Lower is better on every column; – means the model can’t reach that operating point.

ModelFalse cutoffs @ 300 msFalse cutoffs @ 600 msLatency @ 5% cutoffLatency @ 10% cutoff
LiveKit Turn Detector v19.9%4.5%543 ms295 ms
Deepgram Flux12.9%9.9%1151 ms548 ms
ultraVAD27.7%11.9%899 ms663 ms
LiveKit v1-mini27.8%12.1%1070 ms698 ms
Pulse EoT v1 (ours)28.4%8.8%784 ms572 ms
SmartTurn v3.235.2%14.8%1051 ms739 ms
AssemblyAI49.4%14.6%1049 ms713 ms
Soniox–5.5%647 ms512 ms
Cartesia Ink 2––1056 ms911 ms
OpenAI GPT Realtime 2––1143 ms824 ms
VAD baseline55.6%21.7%1600 ms1000 ms

Pulse EoT v1 ranks 3rd of 11 on false cutoffs at 600 ms and 3rd/4th on latency at the 5%/10% budgets - second only to LiveKit v1 among self-hosted models. Against Deepgram Flux specifically: it wins on false cutoffs at 600 ms (8.8% vs 9.9%) and latency at 5% (784 ms vs 1151 ms), and loses on false cutoffs at 300 ms and latency at 10%.

Signal-quality diagnostic (AUC / AP over the full span, independent of any operating-point choice):

ModelAUCAP
LiveKit Turn Detector v10.96860.9414
Pulse EoT v1 (ours)0.96400.9250
Cartesia Ink 20.95490.9091
Deepgram Flux0.94140.8783
Soniox0.90630.8303
AssemblyAI0.89640.7846
LiveKit v1-mini0.89020.8131
ultraVAD0.88440.7966
OpenAI GPT Realtime 20.85840.7436
SmartTurn v3.20.84460.7383

2nd of 10 on raw signal quality. The gap between 2nd on discrimination and mid-pack on false cutoffs at the 300 ms budget is a timing effect, not a signal-quality one - the model knows the turn has ended, it just says so a little late.

Detect rate (correctly identifying a true turn-end at all, independent of timing) is best-in-class: 95.8% at the 5% false-cutoff operating point, vs LiveKit v1’s 91.0%.

Emotion detection

Evaluated on two corpora held out from training:

Test setnUAWAF1
ESD14,00072.4672.4673.98
TESS2,40073.4273.4272.31
Mean (2 held-out sets)—72.9472.9473.14

UA = unweighted accuracy, WA = weighted accuracy. Evaluated across nine emotion classes (anger, happiness, sadness, neutral, excitement, frustration, fear, surprise, disgust); the streaming API surfaces six classes per final - see Response Format.

Streaming introduces essentially no accuracy cost versus the direct (offline) model - per-corpus deltas average within about ±0.3 points, consistent with run-to-run noise rather than degradation.

Gender detection

95% F1, evaluated both intra-corpus and on held-out test sets.


Supported Languages

Pulse 2.0 is English-only.

LanguageCode
Englishen

Requesting any other language value returns a LANGUAGE_NOT_SUPPORTED_BY_MODEL error and closes the socket (see Error Frames). For multilingual or pre-recorded transcription, use standard Pulse (21 streaming + 12 pre-recorded languages).


Authentication

Include your API key in the Authorization header when opening the WebSocket:

Authorization: Bearer SMALLEST_API_KEY

Browsers can’t set headers on a WebSocket - mint a short-lived access token on your server and pass it as the api_key query parameter instead.


Request Parameters

Passed as query parameters on the WebSocket URL, for example:

wss://api.smallest.ai/waves/v1/stt/live?model=pulse-2&language=en&sample_rate=16000&encoding=linear16&emotion_detection=false
ParamTypeDefaultDescription
modelstring—Set to pulse-2 to use Pulse 2.0.
languagestringenOnly en is accepted (or omit it). Any other value returns LANGUAGE_NOT_SUPPORTED_BY_MODEL and closes the socket.
sample_rateinteger—Sample rate of the input audio, in Hz (e.g. 16000).
encodingstring—Audio encoding: linear16, linear32, alaw, mulaw, opus, ogg_opus.
api_keystring—API key for clients that can’t set an Authorization header (e.g. browsers). Use a short-lived access token, not the raw key.
word_timestampsboolfalseAdds a words[] array (word, start, end, confidence) to every final.
sentence_timestampsboolfalseAdds an utterances[] array of sentence-level segments to every final. Requires word_timestamps=true.
diarizeboolfalseAdds per-word speaker / speaker_confidence and a diarization_segments array.
keywordsstring—Boost recognition of custom vocabulary as KEYWORD:INTENSIFIER pairs (e.g. Blackwell:2). See Keyword Boosting for the full syntax.
redact_piiboolfalseRedacts personally identifiable information (names, phone numbers, emails, SSNs).
redact_pciboolfalseRedacts payment-card data (card numbers, CVVs, expiry dates).
itn_normalizeboolfalseInverse text normalization - converts spoken-form numbers, dates, and currency into written form.
formatbooltrueAutomatic punctuation and capitalization. Set false to disable.
vad_eventsboolfalseEmits speech_started / speech_ended message types interleaved with transcription, independent of finalization.
external_session_idstring—Client-supplied id for correlating a session in your own logs. Not echoed back in responses.
emotion_detection NewbooltrueAdds an emotion object to every final. Set false to turn it off.
gender_detection NewbooltrueAdds a gender object to every final. Set false to turn it off.

emotion_detection / gender_detection are the exact param names - emotion=false or gender=false are silently ignored and both fields keep coming back with their defaults.

On pulse-2, finalization is driven by built-in end-of-turn detection, which has no tuning knobs. Pulse’s finalization params eou_timeout_ms, finalize_on_words, and max_words are ignored by default - don’t send them. If you do, they still take effect and compete with end-of-turn detection: for example, max_words=5 splits turns into short fragments and EOT fires less often. If you’re migrating from pulse, remove them from your client. See Finalization and Endpointing for how these work on Pulse.

An unknown model is rejected at the handshake with HTTP 400 INVALID_MODEL.


Control Messages

Sent from the client as JSON text frames on the same WebSocket:

MessageDescription
{"type": "finalize"}Forces the current segment to finalize immediately without closing the connection.
{"type": "close_stream"}Ends the session. Triggers one last final with is_last: true.

Response Format

Every server message is a JSON frame. Interim (is_final: false) messages carry a running transcript string only; finals carry the full set of fields below, subject to which params you set.

FieldTypeWhen presentDescription
typestringAlwaystranscription, speech_started / speech_ended (with vad_events=true), or error.
statusstringAlwayssuccess for valid transcription responses.
session_idstringAlwaysServer-generated id for this connection.
transcriptstringAlwaysPartial or final transcript text for the current segment.
is_finalboolAlwaystrue for a finalized segment, false for an interim/partial one.
is_lastboolAlwaystrue on the last message of the session, sent after close_stream.
language, languagesstring, arrayFinalsLanguage code(s) - always en on Pulse 2.0.
words[]arrayFinals, when word_timestamps=truePer-word word, start, end, confidence; also speaker and speaker_confidence when diarize=true.
utterances[]arrayFinals, when sentence_timestamps=trueSentence-level text, start, end; also speaker when diarize=true.
emotion NewobjectFinals, when emotion_detection is on (default)label, confidence (0–1), and scores for all six classes; scores sum to 1.
emotion.label Newstringwith emotionneutral | happy | angry | sad | fear | disgust
gender NewobjectFinals, when gender_detection is on (default)label, confidence (0–1), and scores for both classes.
gender.label Newstringwith gendermale | female
eot_finalized NewboolFinals closed by end-of-turntrue when the built-in end-of-turn model decided the speaker finished. Absent on the final flushed by close_stream / finalize - that one is marked is_last: true instead.

Real final from prod (words trimmed to two):

{
"type": "transcription",
"status": "success",
"session_id": "e119d986-9ebc-449b-826d-94650fc09a5f",
"transcript": " Good afternoon.",
"is_final": true,
"is_last": false,
"eot_finalized": true,
"language": "en",
"languages": ["en"],
"words": [
{"word": "Good", "start": 1.14, "end": 1.46, "confidence": 0.9517, "speaker": 0, "speaker_confidence": 1},
{"word": "afternoon.", "start": 1.46, "end": 2.1, "confidence": 0.9717, "speaker": 0, "speaker_confidence": 0.875}
],
"emotion": {
"label": "happy", "confidence": 0.5343,
"scores": {"neutral": 0.4515, "happy": 0.5343, "angry": 0.0118, "sad": 0.0023, "fear": 0, "disgust": 0.0001}
},
"gender": {
"label": "male", "confidence": 1,
"scores": {"male": 1, "female": 0}
}
}

Three things to keep in mind when integrating:

  1. Treat emotion and gender as optional. They are left out (not null) when an utterance is too short or classification times out.
  2. Expect fewer, longer finals. On a 25 s speech clip, Pulse 2.0 returned 5 finals, 4 of them eot_finalized, each ending on a sentence or clause.
  3. The last final has no eot_finalized. It comes from closing the stream, marked by is_last: true instead.

Error Frames

{
"type": "error",
"error_code": "LANGUAGE_NOT_SUPPORTED_BY_MODEL",
"message": "model=pulse-2 is English-only (language=en); got language='hi'.",
"model": "pulse-2",
"language": "hi"
}

Error frames always carry type: "error", error_code, and a human-readable message.


API Reference

Endpoint

EndpointMethodUse case
wss://api.smallest.ai/waves/v1/stt/live?model=pulse-2WebSocketStreaming transcription with EOT, emotion, and gender

See Transcribe (Realtime / WebSocket) for the full generated request/response schema, including this page’s params and fields.

There is no batch / pre-recorded Pulse 2.0. For high-volume English batch transcription, use Pulse Pro.


Throughput, Latency & Pricing

Pulse 2.0-specific latency numbers are being finalized. Pulse 2.0 runs on the same streaming infrastructure as Pulse - see Pulse’s Throughput, Latency & Pricing as a reference point in the meantime.

Customer pricing: $0.005 per minute of audio (Standard plan, streaming). Standard plan rate-limit default: 100 concurrent requests. Enterprise tier is unlimited and configurable per-customer.


Use Cases

Strong FitNot a Fit
Voice agents that need to know when the caller finished speakingMultilingual audio - use Pulse instead
Live call analytics that track caller sentiment (emotion)Pre-recorded / batch transcription - use Pulse or Pulse Pro
Demographic routing or reporting on speaker genderIdentity verification or any use of gender as an identity attribute

Limitations

  • English only. Use Pulse for every other language.
  • Emotion and gender are not guaranteed on every final. They’re estimated from the voice of each utterance and can be left out on a very short utterance (under ~0.4 s) or a slow classification pass.
  • Gender is a binary acoustic estimate of the voice, not an identity attribute. It describes how the speaker’s voice sounds to the model, not how the speaker identifies.

Safety & Compliance

Pulse 2.0 must not be used for:

  • Recording or transcribing individuals without their explicit consent
  • Surveillance, stalking, or any form of unauthorized monitoring
  • Any illegal or unethical purposes

Additionally:

  • Usage is monitored for policy compliance
  • Gender output is an acoustic estimate, not an identity signal - do not use it to infer or record a speaker’s gender identity
  • For compliance documentation (GDPR, SOC2, HIPAA), contact support@smallest.ai

FAQ

Pulse 2.0 is English-only and streaming-only. On top of everything Pulse’s streaming mode supports, it adds built-in end-of-turn detection and per-utterance emotion and gender, on by default. For multilingual or batch transcription, use standard Pulse.

Pulse 2.0’s built-in end-of-turn model closes a final when it detects the speaker has finished their turn, rather than on a fixed ~10-word or 800 ms silence window. On a 25 s speech clip this produced 5 finals (4 of them eot_finalized).

Add emotion_detection=false or gender_detection=false to the WebSocket URL - both default to true. Double-check the param name: emotion=false or gender=false are silently ignored and the fields keep coming back.

Both fields are best-effort per-utterance classifications. They’re omitted (not sent as null) when the utterance is too short (under ~0.4 s) for the classifier, or if classification doesn’t complete in time. Treat both fields as optional in your parsing.

eot_finalized: true means the built-in end-of-turn model decided the speaker had finished talking. The very last final of a session comes from the client closing the stream (close_stream) or calling finalize, not from end-of-turn detection - that final is marked with is_last: true instead and has no eot_finalized field.

You shouldn’t. Pulse 2.0 replaces silence- and word-count-based finalization with built-in end-of-turn detection, which has no tuning knobs. eou_timeout_ms, finalize_on_words, and max_words are ignored by default - don’t send them. If you do, they still take effect and compete with end-of-turn detection, so you’ll get shorter, fragmented finals.

No. Pulse 2.0 is streaming-only. For batch transcription, use Pulse or, for English-only high-accuracy batch, Pulse Pro.


Support