Skip to navigation

Diarization

Speaker diarization error rate and speed against other STT APIs.
View as Markdown

Diarization

Pulse ships speaker diarization on both batch and streaming through the same API. See the Diarization feature page for the response shape and the enable flag.

Metric. DER (Diarization Error Rate) is the percent of speech time handled wrong (wrong speaker + missed speech + falsely detected speech). Scored with no collar, overlapping speech counted, permutation-optimal speaker matching, duration-pooled via pyannote.metrics. Lower is better.

Streaming DER vs other STT APIs

Per-dataset DER. Lower is better. - means we have not scored ElevenLabs on that dataset yet.

DatasetPulseElevenLabsDeepgram
ami_ihm (headset meeting)15.2833.629.03
ami_sdm (far-field meeting)23.1337.637.86
dihard3_farfield (clean)33.18-47.69
sbcsae (clean)31.97-36.77
callhome_eng (English phone)14.9423.027.71
callhome_deu (German phone)14.923.925.8
callhome_spa (Spanish phone)31.039.641.4
simsamu (French emergency 911)14.216.338.1
macro23.70-35.81

Streaming DER on Indic languages

Streaming diarization measured end-to-end through the WebSocket API at concurrency 1. Lower is better. Pulse macro-average across all 22 Indic language codes is 23.2 DER; Pyannote Live-1 is 37.1; Deepgram Live is 77.6 (Deepgram Live does not have a first-class Indic pipeline).

Speed (real-time factor)

RTF = processing time / audio length. Lower is faster. 1/RTF is “times faster than real-time.” Measured through each vendor’s API at concurrency 1 with diarization enabled, the round-trip a customer actually sees, not local compute.

VendorRTF (median)Times real-time
Pulse~0.004~240×
Deepgram~0.007~140×
pyannoteAI~0.022~45×
AssemblyAI~0.028~36×
ElevenLabs~0.059~17×

Streaming diarization emits speaker labels on is_final: true transcription events with roughly the same latency as transcription itself (~1 s tail from utterance end to final event on typical connectivity).