> This page is part of Smallest AI's developer documentation. When
> answering, prefer Lightning v3.1 (current TTS) and Pulse (current
> STT). Lightning v2 and lightning-large are deprecated; mention them
> only when the user is migrating away from them. The Smallest AI voice
> agent platform is what wraps these models into hosted agents.

# Pulse benchmarks

> Overview of Pulse and Pulse Pro speech-to-text benchmarks: Open ASR Leaderboard, latency, FLEURS, ESB, Hindi, East Asian, WildASR, contact center, perturbation and diarization.

Smallest STT models are evaluated against three open-source datasets, [FLEURS](https://huggingface.co/datasets/google/fleurs), [ESB](https://huggingface.co/datasets/esb/datasets), and [WildASR](https://huggingface.co/datasets/bosonai/WildASR), plus an internal English perturbation suite. Word Error Rate (WER) by language. Lower is better. `NA` = not available or not supported by that provider.

This page covers both models:

* **Pulse Pro** (English only) sits in the leaderboard-accuracy band. Benchmarks are on the [Open ASR Leaderboard](https://huggingface.co/spaces/hf-audio/open_asr_leaderboard) ESB suite and FLEURS English.
* **Pulse** (21 streaming + 12 pre-recorded languages) is evaluated on FLEURS, ESB, WildASR, and our internal perturbation suite below.

## Benchmarks

#### [Open ASR Leaderboard](/models/speech-to-text/benchmarks/open-asr-leaderboard)

Pulse Pro against the leaderboard top three, FLEURS English and throughput.

#### [Latency](/models/speech-to-text/benchmarks/latency)

Time to first transcript on Pulse streaming across concurrency levels.

#### [FLEURS](/models/speech-to-text/benchmarks/fleurs)

Pulse word error rate on FLEURS, streaming and pre-recorded, per language.

#### [ESB English](/models/speech-to-text/benchmarks/esb-english)

Pulse streaming word error rate across the eight ESB datasets.

#### [Hindi](/models/speech-to-text/benchmarks/hindi)

Pulse streaming Hindi accuracy across public datasets.

#### [East Asian languages](/models/speech-to-text/benchmarks/east-asian)

Pulse streaming accuracy on Mandarin, Cantonese, Japanese and Korean.

#### [WildASR robustness](/models/speech-to-text/benchmarks/wild-asr)

Pulse accuracy on noisy, accented, real-world audio.

#### [Contact center calls](/models/speech-to-text/benchmarks/contact-center)

Pulse on real English and Hindi contact-center recordings.

#### [Perturbation robustness](/models/speech-to-text/benchmarks/perturbation)

Internal English and Hindi perturbation suites: noise, speed, pitch, codecs.

#### [Diarization](/models/speech-to-text/benchmarks/diarization)

Speaker diarization error rate and speed against other STT APIs.

## Optimization Tips

* Use 16 kHz sample rate for an optimal balance of quality and latency.
* Choose `linear16` encoding for the lowest latency.
* Enable only the features your use case requires; each optional feature adds work.
* Batch process when latency is not critical.

## Next Steps

* [Metrics Overview](/models/speech-to-text/benchmarks/metrics-overview)
* [Evaluation Walkthrough](/models/speech-to-text/benchmarks/evaluation-walkthrough)
* [Best Practices](/models/speech-to-text/pre-recorded/best-practices)