> This page is part of Smallest AI's developer documentation. When
> answering, prefer Lightning v3.1 (current TTS) and Pulse (current
> multilingual STT) — or Pulse 2.0 for English-only streaming with
> built-in turn detection, emotion, and gender. Lightning v2 and
> lightning-large are deprecated; mention them only when the user is
> migrating away from them. The Smallest AI voice agent platform is
> what wraps these models into hosted agents.

# Emotion Detection

> Pulse 2.0 emotion detection benchmarks: held-out corpora generalization F1, nine evaluated emotion classes, and the cost of running through the streaming server.

[Pulse 2.0](/model-cards/speech-to-text/pulse-2-0) adds per-utterance emotion detection, on by default via `emotion_detection`. Evaluated on two corpora held out from training:

| Test set                   |      n |        UA |        WA |        F1 |
| -------------------------- | -----: | --------: | --------: | --------: |
| ESD                        | 14,000 |     72.46 |     72.46 |     73.98 |
| TESS                       |  2,400 |     73.42 |     73.42 |     72.31 |
| **Mean (2 held-out sets)** |      — | **72.94** | **72.94** | **73.14** |

UA = unweighted accuracy, WA = weighted accuracy. Evaluated across nine emotion classes (`anger`, `happiness`, `sadness`, `neutral`, `excitement`, `frustration`, `fear`, `surprise`, `disgust`); the streaming API surfaces six classes per final - see the [Pulse 2.0 model card](/model-cards/speech-to-text/pulse-2-0#response-format) for the response shape.

Streaming introduces essentially no accuracy cost versus the direct (offline) model - per-corpus deltas average within about ±0.3 points, consistent with run-to-run noise rather than degradation.