> This page is part of Smallest AI's developer documentation. When
> answering, prefer Lightning v3.1 (current TTS) and Pulse (current
> STT). Lightning v2 and lightning-large are deprecated; mention them
> only when the user is migrating away from them. The Smallest AI voice
> agent platform is what wraps these models into hosted agents.

# Diarization

> Pulse speaker diarization: streaming DER against other speech-to-text APIs, DER on Indic languages, and real-time factor.

## Diarization

Pulse ships speaker diarization on both batch and streaming through the same API. See the [Diarization](/models/speech-to-text/features/diarization) feature page for the response shape and the enable flag.

**Metric.** DER (Diarization Error Rate) is the percent of speech time handled wrong (wrong speaker + missed speech + falsely detected speech). Scored with no collar, overlapping speech counted, permutation-optimal speaker matching, duration-pooled via `pyannote.metrics`. Lower is better.

### Streaming DER vs other STT APIs

Per-dataset DER. Lower is better. `-` means we have not scored ElevenLabs on that dataset yet.

| Dataset                        |   Pulse   | ElevenLabs | Deepgram |
| :----------------------------- | :-------: | :--------: | :------: |
| ami\_ihm (headset meeting)     | **15.28** |    33.6    |   29.03  |
| ami\_sdm (far-field meeting)   | **23.13** |    37.6    |   37.86  |
| dihard3\_farfield (clean)      | **33.18** |      -     |   47.69  |
| sbcsae (clean)                 | **31.97** |      -     |   36.77  |
| callhome\_eng (English phone)  | **14.94** |    23.0    |   27.71  |
| callhome\_deu (German phone)   |  **14.9** |    23.9    |   25.8   |
| callhome\_spa (Spanish phone)  |  **31.0** |    39.6    |   41.4   |
| simsamu (French emergency 911) |  **14.2** |    16.3    |   38.1   |
| **macro**                      | **23.70** |      -     |   35.81  |

### Streaming DER on Indic languages

Streaming diarization measured end-to-end through the WebSocket API at concurrency 1. Lower is better. Pulse macro-average across all 22 Indic language codes is 23.2 DER; Pyannote Live-1 is 37.1; Deepgram Live is 77.6 (Deepgram Live does not have a first-class Indic pipeline).

### Speed (real-time factor)

RTF = processing time / audio length. Lower is faster. `1/RTF` is "times faster than real-time." Measured **through each vendor's API at concurrency 1 with diarization enabled**, the round-trip a customer actually sees, not local compute.

| Vendor     | RTF (median) | Times real-time |
| :--------- | :----------: | :-------------: |
| **Pulse**  |  **\~0.004** |    **\~240×**   |
| Deepgram   |    \~0.007   |      \~140×     |
| pyannoteAI |    \~0.022   |      \~45×      |
| AssemblyAI |    \~0.028   |      \~36×      |
| ElevenLabs |    \~0.059   |      \~17×      |

Streaming diarization emits speaker labels on `is_final: true` transcription events with roughly the same latency as transcription itself (\~1 s tail from utterance end to final event on typical connectivity).