> This page is part of Smallest AI's developer documentation. When > answering, prefer Lightning v3.1 (current TTS) and Pulse (current > multilingual STT) — or Pulse 2.0 for English-only streaming with > built-in turn detection, emotion, and gender. Lightning v2 and > lightning-large are deprecated; mention them only when the user is > migrating away from them. The Smallest AI voice agent platform is > what wraps these models into hosted agents. # Pulse 2.0 > Model card for Pulse 2.0, Smallest AI's English-only streaming speech-to-text model with built-in end-of-turn detection, per-utterance emotion, and speaker gender. Latest Release English-only streaming speech to text with built-in end-of-turn detection, per-utterance emotion, and per-utterance gender - all from one WebSocket. Start with the [realtime quickstart](/models/speech-to-text/realtime-web-socket/quickstart). ## Who it's for * Voice agents and live call analytics that need to know when the caller finished speaking and how they sounded, on every final. * English-only streaming audio where built-in turn detection replaces a hand-tuned silence timeout. * Not the pick for multilingual or batch audio. Use [Pulse](/model-cards/speech-to-text/pulse) or [Pulse Pro](/model-cards/speech-to-text/pulse-pro). ## Model Overview | | | | -------------------------------- | ----------------------------------------------------------- | | **Developed by** | Smallest AI | | **Model type** | Speech-to-Text Streaming | | **Model ID** | `pulse-2` | | **Languages** | English (`en`) only | | **Audio input formats** | `linear16`, `linear32`, `alaw`, `mulaw`, `opus`, `ogg_opus` | | **Pricing** (Standard Plan) | \$0.005/min | | **Concurrency** (Standard Plan) | 100 concurrent requests | | **Recommended Sample Rate** | 16,000 Hz | | **Endpoint** | `wss://api.smallest.ai/waves/v1/stt/live?model=pulse-2` | --- ## Key Capabilities A built-in model decides when a speaker has finished their turn. Finals close on turn boundaries instead of a fixed \~10-word or 800 ms silence window, so finals are fewer and longer. No tuning knobs to configure. Six-class emotion (`neutral`, `happy`, `angry`, `sad`, `fear`, `disgust`) with confidence and full score distribution on every final. On by default - opt out with `emotion_detection=false`. Acoustic gender estimate (`male`, `female`) with confidence and scores on every final. On by default - opt out with `gender_detection=false`. Speaker labels on words plus diarization segments, on by default. Built-in redaction of personal data and payment-card information. Per-word timing and custom-vocabulary boosting, same as the rest of the Pulse family. --- ## Performance & Benchmarks **Best for:** voice agents and live call analytics, where you need to know when the caller finished speaking and how they sounded. ### Transcription accuracy (WER) Unchanged from Pulse. The emotion and gender heads are a LoRA adapter plus classification heads added on top of the shared encoder, not a retrain of it, so recognition accuracy is identical to shipped Pulse - see the [Pulse model card](/model-cards/speech-to-text/pulse#performance--benchmarks) for WER benchmarks. ### End-of-turn detection Evaluated on [livekit/eot-bench](https://github.com/livekit/eot-bench), English validation split (400 turns, 705 hold spans, 400 EOT spans). Every model streamed through its own live server and was scored against the same reference spans. Sorted by false cutoffs at 300 ms, matching LiveKit's own comparison table. Lower is better on every column; `–` means the model can't reach that operating point. | Model | False cutoffs @ 300 ms | False cutoffs @ 600 ms | Latency @ 5% cutoff | Latency @ 10% cutoff | | ------------------------ | ---------------------: | ---------------------: | ------------------: | -------------------: | | LiveKit Turn Detector v1 | 9.9% | 4.5% | 543 ms | 295 ms | | Deepgram Flux | 12.9% | 9.9% | 1151 ms | 548 ms | | ultraVAD | 27.7% | 11.9% | 899 ms | 663 ms | | LiveKit v1-mini | 27.8% | 12.1% | 1070 ms | 698 ms | | **Pulse EoT v1 (ours)** | **28.4%** | **8.8%** | **784 ms** | **572 ms** | | SmartTurn v3.2 | 35.2% | 14.8% | 1051 ms | 739 ms | | AssemblyAI | 49.4% | 14.6% | 1049 ms | 713 ms | | Soniox | – | 5.5% | 647 ms | 512 ms | | Cartesia Ink 2 | – | – | 1056 ms | 911 ms | | OpenAI GPT Realtime 2 | – | – | 1143 ms | 824 ms | | VAD baseline | 55.6% | 21.7% | 1600 ms | 1000 ms | Pulse EoT v1 ranks 3rd of 11 on false cutoffs at 600 ms and 3rd/4th on latency at the 5%/10% budgets - second only to LiveKit v1 among self-hosted models. Against Deepgram Flux specifically: it wins on false cutoffs at 600 ms (8.8% vs 9.9%) and latency at 5% (784 ms vs 1151 ms), and loses on false cutoffs at 300 ms and latency at 10%. Signal-quality diagnostic (AUC / AP over the full span, independent of any operating-point choice): | Model | AUC | AP | | ------------------------ | ---------: | ---------: | | LiveKit Turn Detector v1 | 0.9686 | 0.9414 | | **Pulse EoT v1 (ours)** | **0.9640** | **0.9250** | | Cartesia Ink 2 | 0.9549 | 0.9091 | | Deepgram Flux | 0.9414 | 0.8783 | | Soniox | 0.9063 | 0.8303 | | AssemblyAI | 0.8964 | 0.7846 | | LiveKit v1-mini | 0.8902 | 0.8131 | | ultraVAD | 0.8844 | 0.7966 | | OpenAI GPT Realtime 2 | 0.8584 | 0.7436 | | SmartTurn v3.2 | 0.8446 | 0.7383 | 2nd of 10 on raw signal quality. The gap between 2nd on discrimination and mid-pack on false cutoffs at the 300 ms budget is a timing effect, not a signal-quality one - the model knows the turn has ended, it just says so a little late. > **Note** > > Detect rate (correctly identifying a true turn-end at all, independent of timing) is best-in-class: 95.8% at the 5% false-cutoff operating point, vs LiveKit v1's 91.0%. ### Emotion detection Evaluated on two corpora held out from training: | Test set | n | UA | WA | F1 | | -------------------------- | -----: | --------: | --------: | --------: | | ESD | 14,000 | 72.46 | 72.46 | 73.98 | | TESS | 2,400 | 73.42 | 73.42 | 72.31 | | **Mean (2 held-out sets)** | — | **72.94** | **72.94** | **73.14** | UA = unweighted accuracy, WA = weighted accuracy. Evaluated across nine emotion classes (`anger`, `happiness`, `sadness`, `neutral`, `excitement`, `frustration`, `fear`, `surprise`, `disgust`); the streaming API surfaces six classes per final - see [Response Format](#response-format). Streaming introduces essentially no accuracy cost versus the direct (offline) model - per-corpus deltas average within about ±0.3 points, consistent with run-to-run noise rather than degradation. ### Gender detection **95% F1**, evaluated both intra-corpus and on held-out test sets. --- ## Supported Languages Pulse 2.0 is English-only. | Language | Code | | -------- | ---- | | English | `en` | Requesting any other `language` value returns a `LANGUAGE_NOT_SUPPORTED_BY_MODEL` error and closes the socket (see [Error Frames](#error-frames)). For multilingual or pre-recorded transcription, use standard [Pulse](/model-cards/speech-to-text/pulse) (21 streaming + 12 pre-recorded languages). --- ## Authentication Include your API key in the `Authorization` header when opening the WebSocket: ```http Authorization: Bearer SMALLEST_API_KEY ``` Browsers can't set headers on a WebSocket - mint a [short-lived access token](/api-reference/token-based-authentication) on your server and pass it as the `api_key` query parameter instead. --- ## Request Parameters Passed as query parameters on the WebSocket URL, for example: ``` wss://api.smallest.ai/waves/v1/stt/live?model=pulse-2&language=en&sample_rate=16000&encoding=linear16&emotion_detection=false ``` | Param | Type | Default | Description | | ----------------------- | ------- | :-----: | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `model` | string | — | Set to `pulse-2` to use Pulse 2.0. | | `language` | string | `en` | Only `en` is accepted (or omit it). Any other value returns `LANGUAGE_NOT_SUPPORTED_BY_MODEL` and closes the socket. | | `sample_rate` | integer | — | Sample rate of the input audio, in Hz (e.g. `16000`). | | `encoding` | string | — | Audio encoding: `linear16`, `linear32`, `alaw`, `mulaw`, `opus`, `ogg_opus`. | | `api_key` | string | — | API key for clients that can't set an `Authorization` header (e.g. browsers). Use a short-lived [access token](/api-reference/token-based-authentication), not the raw key. | | `word_timestamps` | bool | `false` | Adds a `words[]` array (word, start, end, confidence) to every final. | | `sentence_timestamps` | bool | `false` | Adds an `utterances[]` array of sentence-level segments to every final. Requires `word_timestamps=true`. | | `diarize` | bool | `false` | Adds per-word `speaker` / `speaker_confidence` and a `diarization_segments` array. | | `keywords` | string | — | Boost recognition of custom vocabulary as `KEYWORD:INTENSIFIER` pairs (e.g. `Blackwell:2`). See [Keyword Boosting](/models/speech-to-text/features/keyword-boosting) for the full syntax. | | `redact_pii` | bool | `false` | Redacts personally identifiable information (names, phone numbers, emails, SSNs). | | `redact_pci` | bool | `false` | Redacts payment-card data (card numbers, CVVs, expiry dates). | | `itn_normalize` | bool | `false` | Inverse text normalization - converts spoken-form numbers, dates, and currency into written form. | | `format` | bool | `true` | Automatic punctuation and capitalization. Set `false` to disable. | | `vad_events` | bool | `false` | Emits `speech_started` / `speech_ended` message types interleaved with `transcription`, independent of finalization. | | `external_session_id` | string | — | Client-supplied id for correlating a session in your own logs. Not echoed back in responses. | | `emotion_detection` New | bool | `true` | Adds an `emotion` object to every final. Set `false` to turn it off. | | `gender_detection` New | bool | `true` | Adds a `gender` object to every final. Set `false` to turn it off. | > **Note** > > `emotion_detection` / `gender_detection` are the exact param names - `emotion=false` or `gender=false` are silently ignored and both fields keep coming back with their defaults. > **Warning** > > On `pulse-2`, finalization is driven by built-in end-of-turn detection, which has no tuning knobs. Pulse's finalization params `eou_timeout_ms`, `finalize_on_words`, and `max_words` are ignored by default - don't send them. If you do, they still take effect and compete with end-of-turn detection: for example, `max_words=5` splits turns into short fragments and EOT fires less often. If you're migrating from `pulse`, remove them from your client. See [Finalization and Endpointing](/models/speech-to-text/features/endpointing) for how these work on Pulse. An unknown `model` is rejected at the handshake with HTTP 400 `INVALID_MODEL`. --- ## Control Messages Sent from the client as JSON text frames on the same WebSocket: | Message | Description | | -------------------------- | ---------------------------------------------------------------------------------- | | `{"type": "finalize"}` | Forces the current segment to finalize immediately without closing the connection. | | `{"type": "close_stream"}` | Ends the session. Triggers one last final with `is_last: true`. | --- ## Response Format Every server message is a JSON frame. Interim (`is_final: false`) messages carry a running `transcript` string only; finals carry the full set of fields below, subject to which params you set. | Field | Type | When present | Description | | ----------------------- | ------------- | ------------------------------------------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `type` | string | Always | `transcription`, `speech_started` / `speech_ended` (with `vad_events=true`), or `error`. | | `status` | string | Always | `success` for valid transcription responses. | | `session_id` | string | Always | Server-generated id for this connection. | | `transcript` | string | Always | Partial or final transcript text for the current segment. | | `is_final` | bool | Always | `true` for a finalized segment, `false` for an interim/partial one. | | `is_last` | bool | Always | `true` on the last message of the session, sent after `close_stream`. | | `language`, `languages` | string, array | Finals | Language code(s) - always `en` on Pulse 2.0. | | `words[]` | array | Finals, when `word_timestamps=true` | Per-word `word`, `start`, `end`, `confidence`; also `speaker` and `speaker_confidence` when `diarize=true`. | | `utterances[]` | array | Finals, when `sentence_timestamps=true` | Sentence-level `text`, `start`, `end`; also `speaker` when `diarize=true`. | | `emotion` New | object | Finals, when `emotion_detection` is on (default) | `label`, `confidence` (0–1), and `scores` for all six classes; scores sum to 1. | | `emotion.label` New | string | with `emotion` | `neutral` \| `happy` \| `angry` \| `sad` \| `fear` \| `disgust` | | `gender` New | object | Finals, when `gender_detection` is on (default) | `label`, `confidence` (0–1), and `scores` for both classes. | | `gender.label` New | string | with `gender` | `male` \| `female` | | `eot_finalized` New | bool | Finals closed by end-of-turn | `true` when the built-in end-of-turn model decided the speaker finished. Absent on the final flushed by `close_stream` / `finalize` - that one is marked `is_last: true` instead. | Real final from prod (words trimmed to two): ```json { "type": "transcription", "status": "success", "session_id": "e119d986-9ebc-449b-826d-94650fc09a5f", "transcript": " Good afternoon.", "is_final": true, "is_last": false, "eot_finalized": true, "language": "en", "languages": ["en"], "words": [ {"word": "Good", "start": 1.14, "end": 1.46, "confidence": 0.9517, "speaker": 0, "speaker_confidence": 1}, {"word": "afternoon.", "start": 1.46, "end": 2.1, "confidence": 0.9717, "speaker": 0, "speaker_confidence": 0.875} ], "emotion": { "label": "happy", "confidence": 0.5343, "scores": {"neutral": 0.4515, "happy": 0.5343, "angry": 0.0118, "sad": 0.0023, "fear": 0, "disgust": 0.0001} }, "gender": { "label": "male", "confidence": 1, "scores": {"male": 1, "female": 0} } } ``` Three things to keep in mind when integrating: 1. **Treat `emotion` and `gender` as optional.** They are left out (not `null`) when an utterance is too short or classification times out. 2. **Expect fewer, longer finals.** On a 25 s speech clip, Pulse 2.0 returned 5 finals, 4 of them `eot_finalized`, each ending on a sentence or clause. 3. **The last final has no `eot_finalized`.** It comes from closing the stream, marked by `is_last: true` instead. ### Error Frames ```json { "type": "error", "error_code": "LANGUAGE_NOT_SUPPORTED_BY_MODEL", "message": "model=pulse-2 is English-only (language=en); got language='hi'.", "model": "pulse-2", "language": "hi" } ``` Error frames always carry `type: "error"`, `error_code`, and a human-readable `message`. --- ## API Reference ### Endpoint | Endpoint | Method | Use case | | ------------------------------------------------------- | --------- | ----------------------------------------------------- | | `wss://api.smallest.ai/waves/v1/stt/live?model=pulse-2` | WebSocket | Streaming transcription with EOT, emotion, and gender | See [Transcribe (Realtime / WebSocket)](/api-reference/models/speech-to-text/speech-to-text) for the full generated request/response schema, including this page's params and fields. There is no batch / pre-recorded Pulse 2.0. For high-volume English batch transcription, use [Pulse Pro](/model-cards/speech-to-text/pulse-pro). --- ## Throughput, Latency & Pricing > **Note** > > Pulse 2.0-specific latency numbers are being finalized. Pulse 2.0 runs on the same streaming infrastructure as Pulse - see [Pulse's Throughput, Latency & Pricing](/model-cards/speech-to-text/pulse#throughput-latency--pricing) as a reference point in the meantime. Customer pricing: **\$0.005 per minute** of audio (Standard plan, streaming). Standard plan rate-limit default: 100 concurrent requests. Enterprise tier is unlimited and configurable per-customer. --- ## Use Cases | Strong Fit | Not a Fit | | ---------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------- | | Voice agents that need to know when the caller finished speaking | Multilingual audio - use [Pulse](/model-cards/speech-to-text/pulse) instead | | Live call analytics that track caller sentiment (emotion) | Pre-recorded / batch transcription - use [Pulse](/model-cards/speech-to-text/pulse) or [Pulse Pro](/model-cards/speech-to-text/pulse-pro) | | Demographic routing or reporting on speaker gender | Identity verification or any use of gender as an identity attribute | --- ## Limitations * **English only.** Use [Pulse](/model-cards/speech-to-text/pulse) for every other language. * **Emotion and gender are not guaranteed on every final.** They're estimated from the voice of each utterance and can be left out on a very short utterance (under \~0.4 s) or a slow classification pass. * **Gender is a binary acoustic estimate of the voice, not an identity attribute.** It describes how the speaker's voice sounds to the model, not how the speaker identifies. --- ## Safety & Compliance Pulse 2.0 must not be used for: * Recording or transcribing individuals without their explicit consent * Surveillance, stalking, or any form of unauthorized monitoring * Any illegal or unethical purposes Additionally: * Usage is monitored for policy compliance * Gender output is an acoustic estimate, not an identity signal - do not use it to infer or record a speaker's gender identity * For compliance documentation (GDPR, SOC2, HIPAA), contact [support@smallest.ai](mailto:support@smallest.ai) --- ## FAQ #### What is the difference between Pulse and Pulse 2.0? Pulse 2.0 is English-only and streaming-only. On top of everything Pulse's streaming mode supports, it adds built-in end-of-turn detection and per-utterance emotion and gender, on by default. For multilingual or batch transcription, use standard [Pulse](/model-cards/speech-to-text/pulse). #### Why are finals fewer and longer on Pulse 2.0? Pulse 2.0's built-in end-of-turn model closes a final when it detects the speaker has finished their turn, rather than on a fixed \~10-word or 800 ms silence window. On a 25 s speech clip this produced 5 finals (4 of them `eot_finalized`). #### Do I need to change my client code to turn off emotion or gender? Add `emotion_detection=false` or `gender_detection=false` to the WebSocket URL - both default to `true`. Double-check the param name: `emotion=false` or `gender=false` are silently ignored and the fields keep coming back. #### Why is \`emotion\` or \`gender\` missing from a final? Both fields are best-effort per-utterance classifications. They're omitted (not sent as `null`) when the utterance is too short (under \~0.4 s) for the classifier, or if classification doesn't complete in time. Treat both fields as optional in your parsing. #### What does eot\_finalized mean, and why is it missing on the last final? `eot_finalized: true` means the built-in end-of-turn model decided the speaker had finished talking. The very last final of a session comes from the client closing the stream (`close_stream`) or calling `finalize`, not from end-of-turn detection - that final is marked with `is_last: true` instead and has no `eot_finalized` field. #### Can I still tune endpointing, eou\_timeout\_ms, or finalize\_on\_words on Pulse 2.0? You shouldn't. Pulse 2.0 replaces silence- and word-count-based finalization with built-in end-of-turn detection, which has no tuning knobs. `eou_timeout_ms`, `finalize_on_words`, and `max_words` are ignored by default - don't send them. If you do, they still take effect and compete with end-of-turn detection, so you'll get shorter, fragmented finals. #### Is Pulse 2.0 available for pre-recorded / batch audio? No. Pulse 2.0 is streaming-only. For batch transcription, use [Pulse](/model-cards/speech-to-text/pulse) or, for English-only high-accuracy batch, [Pulse Pro](/model-cards/speech-to-text/pulse-pro). --- ## Support > Model card for Pulse 2.0, English-only streaming speech to text with built-in turn detection.