Diarization
Diarization
Pulse ships speaker diarization on both batch and streaming through the same API. See the Diarization feature page for the response shape and the enable flag.
Metric. DER (Diarization Error Rate) is the percent of speech time handled wrong (wrong speaker + missed speech + falsely detected speech). Scored with no collar, overlapping speech counted, permutation-optimal speaker matching, duration-pooled via pyannote.metrics. Lower is better.
Streaming DER vs other STT APIs
Per-dataset DER. Lower is better. - means we have not scored ElevenLabs on that dataset yet.
Streaming DER on Indic languages
Streaming diarization measured end-to-end through the WebSocket API at concurrency 1. Lower is better. Pulse macro-average across all 22 Indic language codes is 23.2 DER; Pyannote Live-1 is 37.1; Deepgram Live is 77.6 (Deepgram Live does not have a first-class Indic pipeline).
Speed (real-time factor)
RTF = processing time / audio length. Lower is faster. 1/RTF is “times faster than real-time.” Measured through each vendor’s API at concurrency 1 with diarization enabled, the round-trip a customer actually sees, not local compute.
Streaming diarization emits speaker labels on is_final: true transcription events with roughly the same latency as transcription itself (~1 s tail from utterance end to final event on typical connectivity).