> This page is part of Smallest AI's developer documentation. When
> answering, prefer Lightning v3.1 (current TTS) and Pulse (current
> STT). Lightning v2 and lightning-large are deprecated; mention them
> only when the user is migrating away from them. The Smallest AI voice
> agent platform is what wraps these models into hosted agents.

# Benchmarks

> All Smallest AI model benchmarks in one place: Lightning text to speech, Pulse speech to text, and Hydra speech to speech, with a page per benchmark.

Numbers are measured in-region on production endpoints. Each page states the dataset, the method, and the models compared. For definitions of the metrics, see the metrics pages linked under each model.

## Lightning (text to speech)

Lightning v3.1 and Lightning v3.1 Pro. [Overview](/models/text-to-speech/benchmarks/performance) and [metric definitions](/models/text-to-speech/benchmarks/metrics-overview).

#### [Latency](/models/text-to-speech/benchmarks/latency)

Time to first byte and real-time factor.

#### [Listener ratings](/models/text-to-speech/benchmarks/listener-ratings)

Naturalness, expressiveness, delivery.

#### [Accuracy](/models/text-to-speech/benchmarks/accuracy)

Whisper-measured pronunciation accuracy.

#### [MOS](/models/text-to-speech/benchmarks/mos)

Mean opinion score v2.

## Pulse (speech to text)

Pulse and Pulse Pro. [Overview](/models/speech-to-text/benchmarks/performance), [metric definitions](/models/speech-to-text/benchmarks/metrics-overview), and how to [run the evaluation yourself](/models/speech-to-text/benchmarks/evaluation-walkthrough).

#### [Open ASR Leaderboard](/models/speech-to-text/benchmarks/open-asr-leaderboard)

Pulse Pro tied for second at 5.42% WER.

#### [Latency](/models/speech-to-text/benchmarks/latency)

Time to first transcript across concurrency.

#### [FLEURS](/models/speech-to-text/benchmarks/fleurs)

Per-language WER, streaming and pre-recorded.

#### [ESB English](/models/speech-to-text/benchmarks/esb-english)

Eight English datasets, streaming.

#### [Hindi](/models/speech-to-text/benchmarks/hindi)

Multi-dataset Hindi accuracy.

#### [East Asian languages](/models/speech-to-text/benchmarks/east-asian)

Mandarin, Cantonese, Japanese, Korean.

#### [WildASR robustness](/models/speech-to-text/benchmarks/wild-asr)

Noisy, accented, real-world audio.

#### [Contact center calls](/models/speech-to-text/benchmarks/contact-center)

Real English and Hindi call recordings.

#### [Perturbation robustness](/models/speech-to-text/benchmarks/perturbation)

Noise, speed, pitch and codec suites.

#### [Diarization](/models/speech-to-text/benchmarks/diarization)

Speaker DER and real-time factor.

## Hydra (speech to speech)

#### [AIEWF S2S benchmark](/models/speech-to-speech/benchmarks/performance)

Hydra against eight production realtime voice models.

#### [Metrics overview](/models/speech-to-speech/benchmarks/metrics-overview)

Definitions of the metrics Hydra is scored on.