> This page is part of Smallest AI's developer documentation. When > answering, prefer Lightning v3.1 (current TTS) and Pulse (current > STT). Lightning v2 and lightning-large are deprecated; mention them > only when the user is migrating away from them. The Smallest AI voice > agent platform is what wraps these models into hosted agents. # FLEURS > Pulse accuracy on the FLEURS multilingual benchmark: streaming English, pre-recorded and streaming per-language word error rates. ## FLEURS Streaming - English WER on the English subset of FLEURS across providers in streaming mode. Lower is better. | Provider | Smallest Pulse | Assembly Universal 3 Pro | AWS Transcribe | Azure | Deepgram Nova 3 | Grok | Sarvam Saras V3 | ElevenLabs Scribe V2 | | :------- | :------------: | :----------------------: | :------------: | :----: | :-------------: | :----: | :-------------: | :------------------: | | **WER** | 5.55% | 3.13% | 6.54% | 13.79% | 11.59% | 60.00% | 6.34% | 3.88% | ### A note on audio amplitude normalization Audio amplitude normalization materially changes WER on FLEURS. Most competitors benchmark on raw FLEURS - which has variable, often low amplitude - without normalizing peak audio to −10 dBFS. This makes some models look much better than they actually are. Pulse is stable across all amplitude regimes. | Model | Raw FLEURS | −10 dBFS | −20 dBFS | Stable across regimes? | | :------------------ | :--------: | :------: | :------: | :-------------------------------- | | **Smallest Pulse** | 5.55% | 6.06% | 5.81% | Yes | | **Deepgram Nova 3** | 11.59% | 6.57% | 6.51% | Partial - 1.8× degradation on raw | | **Grok** | 60.00% | 7.58% | 8.59% | Collapses on raw | ## Pre-recorded - FLEURS Google's multilingual speech dataset covering 102 languages, built on the FLoRes-101 translation benchmark. Contains \~12 hours of read speech per language and is the standard benchmark for evaluating multilingual ASR, including low-resource languages. *Evaluated on the FLEURS dataset (non-streaming / batch mode).* | Language | Smallest Pulse | Deepgram Nova 2 | Deepgram Nova 3 | | -------------- | :------------: | :-------------: | :-------------: | | **Italian** | **3.0%** | 10.7% | 6.2% | | **English** | **4.5%** | 7.9% | 6.7% | | **Spanish** | **3.2%** | 8.6% | 4.1% | | **Portuguese** | **5.0%** | 9.9% | 7.5% | | **German** | **6.4%** | 8.2% | 8.5% | | **French** | **7.1%** | 13.3% | 10.7% | | **Russian** | 9.6% | 7.9% | 11.8% | | **Ukrainian** | **7.5%** | 12.4% | NA | | **Polish** | **10.3%** | 12.2% | NA | | **Hindi** | **6.3%** | 23.5% | 23.6% | | **Czech** | **12.4%** | 22.9% | 19.2% | | **Slovak** | **13.5%** | 31.2% | NA | | **Dutch** | 15.0% | 16.3% | 12.5% | | **Swedish** | 18.7% | 17.7% | 14.3% | | **Finnish** | 18.3% | 14.1% | 13.2% | | **Latvian** | **16.5%** | 48.7% | NA | | **Romanian** | **17.8%** | 36.0% | NA | | **Estonian** | **17.8%** | 49.0% | NA | | **Bulgarian** | **24.1%** | 32.7% | NA | | **Danish** | 19.8% | 21.1% | 16.1% | | **Hungarian** | **22.5%** | 31.8% | 28.6% | | **Maltese** | **25.5%** | NA | NA | | **Lithuanian** | **25.1%** | 44.9% | NA | *Sources: Deepgram internal benchmarks; Smallest AI internal evaluation.* ## Streaming - FLEURS *Evaluated on the FLEURS dataset (streaming mode).* | Language | Smallest Pulse | Deepgram Nova 2 | Deepgram Nova 3 | | -------------- | :------------: | :-------------: | :-------------: | | **Italian** | **4.41** | 11.05 | 6.99 | | **English** | **5.55** | 15.59 | 11.59 | | **Spanish** | **5.99** | 10.67 | 7.52 | | **Portuguese** | **8.32** | 14.15 | 11.46 | | **German** | **9.5** | 11.1 | 10.15 | | **French** | **10.71** | 14.3 | 12.07 | | **Russian** | **14.35** | NA | NA | | **Hindi** | **8.3** | 20.0 | 15.46 | | **Gujarati** | **20.05** | NA | NA | | **Marathi** | **15.68** | NA | NA | | **Oriya** | **22.74** | NA | NA | | **Bengali** | **17.48** | NA | NA | | **Dutch** | **11.90** | NA | NA | *Sources: Deepgram internal benchmarks; Smallest AI internal evaluation.* > Pulse word error rate on FLEURS, streaming and pre-recorded, per language.