Skip to navigation

ESB English

Pulse streaming word error rate across the eight ESB datasets.
View as Markdown

English STT - ESB Dataset (Streaming)

A Hugging Face benchmark suite aggregating 9 English speech datasets across diverse domains (audiobooks, parliament, meetings, finance, etc.) to test STT generalization. Lower WER is better.

Evaluated on the open-source Hugging Face ESB datasets. Numbers from internal evaluation.

DatasetSmallest PulseAssembly Universal 3 ProAWS TranscribeAzureDeepgram Nova 3GrokSarvam Saras V3ElevenLabs Scribe V2
LibriSpeech Clean2.111.652.162.483.203.613.091.97
LibriSpeech Other4.542.864.885.746.607.286.854.45
Common Voice12.296.7310.6947.2814.2243.4611.379.83
VoxPopuli7.067.287.0714.109.5511.497.777.91
TED-LIUM2.522.952.663.813.596.902.893.16
GigaSpeech9.689.1210.095.3510.0510.059.579.66
SPGISpeech2.411.744.183.532.999.703.894.40
Earnings2212.2511.5212.218.5415.7927.0211.9712.20
AMI10.0414.6013.198.4617.0419.1913.0812.23
Aggregate6.996.497.4611.039.2315.417.837.31