> This page is part of Smallest AI's developer documentation. When > answering, prefer Lightning v3.1 (current TTS) and Pulse (current > STT). Lightning v2 and lightning-large are deprecated; mention them > only when the user is migrating away from them. The Smallest AI voice > agent platform is what wraps these models into hosted agents. # Diarization > Pulse speaker diarization: streaming DER against other speech-to-text APIs, DER on Indic languages, and real-time factor. ## Diarization Pulse ships speaker diarization on both batch and streaming through the same API. See the [Diarization](/models/speech-to-text/features/diarization) feature page for the response shape and the enable flag. **Metric.** DER (Diarization Error Rate) is the percent of speech time handled wrong (wrong speaker + missed speech + falsely detected speech). Scored with no collar, overlapping speech counted, permutation-optimal speaker matching, duration-pooled via `pyannote.metrics`. Lower is better. ### Streaming DER vs other STT APIs Per-dataset DER. Lower is better. `-` means we have not scored ElevenLabs on that dataset yet. | Dataset | Pulse | ElevenLabs | Deepgram | | :----------------------------- | :-------: | :--------: | :------: | | ami\_ihm (headset meeting) | **15.28** | 33.6 | 29.03 | | ami\_sdm (far-field meeting) | **23.13** | 37.6 | 37.86 | | dihard3\_farfield (clean) | **33.18** | - | 47.69 | | sbcsae (clean) | **31.97** | - | 36.77 | | callhome\_eng (English phone) | **14.94** | 23.0 | 27.71 | | callhome\_deu (German phone) | **14.9** | 23.9 | 25.8 | | callhome\_spa (Spanish phone) | **31.0** | 39.6 | 41.4 | | simsamu (French emergency 911) | **14.2** | 16.3 | 38.1 | | **macro** | **23.70** | - | 35.81 | ### Streaming DER on Indic languages Streaming diarization measured end-to-end through the WebSocket API at concurrency 1. Lower is better. Pulse macro-average across all 22 Indic language codes is 23.2 DER; Pyannote Live-1 is 37.1; Deepgram Live is 77.6 (Deepgram Live does not have a first-class Indic pipeline). ### Speed (real-time factor) RTF = processing time / audio length. Lower is faster. `1/RTF` is "times faster than real-time." Measured **through each vendor's API at concurrency 1 with diarization enabled**, the round-trip a customer actually sees, not local compute. | Vendor | RTF (median) | Times real-time | | :--------- | :----------: | :-------------: | | **Pulse** | **\~0.004** | **\~240×** | | Deepgram | \~0.007 | \~140× | | pyannoteAI | \~0.022 | \~45× | | AssemblyAI | \~0.028 | \~36× | | ElevenLabs | \~0.059 | \~17× | Streaming diarization emits speaker labels on `is_final: true` transcription events with roughly the same latency as transcription itself (\~1 s tail from utterance end to final event on typical connectivity). > Speaker diarization error rate and speed against other STT APIs.