> This page is part of Smallest AI's developer documentation. When
> answering, prefer Lightning v3.1 (current TTS) and Pulse (current
> STT). Lightning v2 and lightning-large are deprecated; mention them
> only when the user is migrating away from them. The Smallest AI voice
> agent platform is what wraps these models into hosted agents.

# Speaker diarization

> Attach zero-indexed speaker labels to each word and utterance in a transcript on Pulse batch and streaming.

Pre-Recorded

Real-Time

Diarization labels each word and utterance with an integer speaker id in the same response you already get for transcription. One API call, no separate diarization vendor, no reconciliation of two timelines.

Available on Pulse in both batch (`POST /waves/v1/stt/`) and streaming (`WSS /waves/v1/stt/live`).

Language-agnostic: works across every language Pulse supports.

## Enable

Add `diarize=true`. Pair with `word_timestamps=true` if you want per-word speaker labels (per-utterance labels come regardless).

### Batch (Pre-Recorded)

```bash
curl -sL -o audio.wav "https://github.com/smallest-inc/cookbook/raw/main/speech-to-text/getting-started/samples/audio.wav"

curl --request POST \
  --url "https://api.smallest.ai/waves/v1/stt/?model=pulse&language=en&diarize=true&word_timestamps=true" \
  --header "Authorization: Bearer $SMALLEST_API_KEY" \
  --header "Content-Type: audio/wav" \
  --data-binary "@audio.wav"
```

### Streaming (Real-Time WebSocket)

```javascript
const url = new URL("wss://api.smallest.ai/waves/v1/stt/live?model=pulse");
url.searchParams.append("language", "en");
url.searchParams.append("encoding", "linear16");
url.searchParams.append("sample_rate", "16000");
url.searchParams.append("diarize", "true");
url.searchParams.append("word_timestamps", "true");

const ws = new WebSocket(url.toString(), {
  headers: { Authorization: `Bearer ${API_KEY}` },
});
```

## Response shape

The response carries two arrays that both gain speaker labels when `diarize=true`:

* `words[]`. Per-word entries. Present when `word_timestamps=true` (batch) or on final events (streaming). Each entry gains `speaker` (integer) and `speaker_confidence` (0.0 to 1.0) when diarization is on.
* `utterances[]`. Sentence-level entries. Present by default on batch, and on streaming with `sentence_timestamps=true`. Each entry gains `speaker` (integer) when diarization is on.

Speaker labels are **integers, zero-indexed** (`0`, `1`, `2`, ...) on both batch and streaming. Labels are stable within a single response or session, not across separate requests. Speaker `0` in one call is not necessarily the same physical speaker as speaker `0` in the next.

### Batch response

```json
{
  "status": "success",
  "transcription": "Agent: Hello, how can I help. Customer: I need to reset my password.",
  "words": [
    { "start": 0.24, "end": 0.40, "word": "Hello,",   "confidence": 0.998, "speaker": 0, "speaker_confidence": 1.00 },
    { "start": 0.40, "end": 0.72, "word": "how",      "confidence": 0.996, "speaker": 0, "speaker_confidence": 1.00 },
    { "start": 0.72, "end": 1.10, "word": "can",      "confidence": 0.999, "speaker": 0, "speaker_confidence": 1.00 },
    { "start": 1.20, "end": 1.60, "word": "I",        "confidence": 0.999, "speaker": 1, "speaker_confidence": 0.92 },
    { "start": 1.60, "end": 1.92, "word": "need",     "confidence": 0.997, "speaker": 1, "speaker_confidence": 0.94 }
  ],
  "utterances": [
    { "start": 0.24, "end": 1.10, "speaker": 0, "text": "Hello, how can I help." },
    { "start": 1.20, "end": 3.44, "speaker": 1, "text": "I need to reset my password." }
  ]
}
```

### Streaming response

Streaming emits **`transcription` events**. Only events with `is_final: true` carry the `words[]` array; interim (partial) events send a running `transcript` string only.

```json
// interim event (no words, no speaker attribution yet)
{ "type": "transcription", "transcript": "Hello, how can I", "is_final": false, "is_last": false }

// final event (words[] with speaker + speaker_confidence)
{
  "type": "transcription",
  "transcript": "Hello, how can I help.",
  "is_final": true,
  "is_last": false,
  "words": [
    { "word": "Hello,", "start": 0.32, "end": 0.80, "confidence": 0.964, "speaker": 0, "speaker_confidence": 0.87 },
    { "word": "how",    "start": 0.80, "end": 0.88, "confidence": 0.931, "speaker": 0, "speaker_confidence": 1.00 },
    { "word": "can",    "start": 1.28, "end": 1.52, "confidence": 0.959, "speaker": 0, "speaker_confidence": 1.00 },
    { "word": "I",      "start": 1.52, "end": 2.00, "confidence": 0.980, "speaker": 0, "speaker_confidence": 1.00 },
    { "word": "help.",  "start": 2.40, "end": 2.96, "confidence": 0.943, "speaker": 0, "speaker_confidence": 1.00 }
  ]
}
```

Practical consequences of the interim/final split:

* **Do not render speaker attribution off interim events.** They are transcript-string only. Speaker labels are correct only from `is_final: true` events.
* If you are driving live captions and want a running speaker chip on-screen, hold the current speaker from the last `is_final` event and update on the next one.

## Response fields

| Field                         | Type             | When present                                                               | Notes                                               |
| ----------------------------- | ---------------- | -------------------------------------------------------------------------- | --------------------------------------------------- |
| `words[i].speaker`            | integer          | `diarize=true` (batch: always; streaming: only on `is_final=true`)         | Zero-indexed. Stable within one response/session.   |
| `words[i].speaker_confidence` | number (0.0–1.0) | `diarize=true` (same visibility as `words[i].speaker`)                     | Confidence that the speaker attribution is correct. |
| `utterances[i].speaker`       | integer          | `diarize=true` (batch: always; streaming: with `sentence_timestamps=true`) | Zero-indexed. Utterances are turn-length segments.  |

Every other field on `words` (`word`, `start`, `end`, `confidence`) is present regardless of `diarize`. Same for `utterances` (`text`, `start`, `end`).

## Gotchas

* **Pulse Pro does not diarize.** `POST /waves/v1/stt/?model=pulse-pro&diarize=true` returns a transcript-only response with no `speaker` fields on `words[]` and no `utterances[]`. Route diarization traffic to `?model=pulse`.
* **Interim events on streaming carry no speaker.** Only finals do. See the note above.
* **Labels are not identity.** `speaker: 0` in call A and `speaker: 0` in call B are unrelated. If you need to persist "this is Alice" across calls, do the mapping yourself (voice-print embedding, dashboard tagging, etc).
* **`word_timestamps=true` is required for per-word speakers on batch.** Without it, `words[]` is empty and only `utterances[]` carries speaker labels.