Skip to navigation

Accuracy

Pronunciation accuracy measured with Whisper jiwer and Whisper LLM.
View as Markdown

Accuracy

Mixed direction - WER, CER, Hallucination, and Deletion are lower is better; Pronunciation % is higher is better.

Whisper jiwer

MetricDirectionLightning v3.1Lightning v3.1 ProGPT-4o-miniElevenLabs Turbo v2.5ElevenLabs Multilingual v2Sonic-3Gemini 2.5 ProGemini 2.5 FlashMAI-Voice-1Inworld 1.5S2 Pro
WERlower1.57%1.36%1.26%1.35%1.33%1.43%1.26%1.37%1.25%1.10%2.83%
CERlower0.67%0.40%0.52%0.60%0.54%0.59%0.62%0.61%0.50%0.47%1.16%
Hallucinationlower0.03%0.00%0.07%0.08%0.01%0.06%0.04%0.01%0.06%0.00%0.22%
DeletionlowerNA0.00%0.14%0.17%0.18%0.16%0.24%0.18%0.15%0.12%0.33%
Pronunciation %
Whisper jiwer
higher98.61%98.68%98.94%98.90%98.87%98.79%99.02%98.82%98.95%99.02%97.72%

Whisper LLM (Pro evaluation only)

LLM-judged Whisper transcripts, applied during the Pro benchmark run. The follow-on LLM normalizes punctuation, casing, and Whisper’s own transcription noise - typically reducing false-positive errors compared to jiwer. Standard Lightning v3.1 was not evaluated with this methodology.

MetricDirectionLightning v3.1 ProGPT-4o-miniElevenLabs Turbo v2.5ElevenLabs Multilingual v2Sonic-3Gemini 2.5 ProGemini 2.5 FlashMAI-Voice-1Inworld 1.5S2 Pro
WERlower0.96%0.82%0.72%0.57%0.88%0.70%0.72%0.60%0.55%2.15%
CERlower0.34%0.30%0.28%0.21%0.30%0.35%0.33%0.23%0.18%1.03%
Hallucinationlower0.00%0.07%0.07%0.00%0.02%0.02%0.01%0.03%0.00%0.10%
Pronunciation %
Whisper LLM
higher99.04%99.25%99.35%99.43%99.14%99.32%99.29%99.43%99.45%97.95%
  • WER (Word Error Rate) - Percentage of words in the transcript that differ from the reference; measures how faithfully the TTS renders the input text.
  • CER (Character Error Rate) - Like WER but at the character level.
  • Hallucination - Words or sounds the TTS generates that have no basis in the input text. Insertions, substitutions, or fabricated content.
  • Deletion - Words from the reference text that the TTS dropped entirely.
  • Pronunciation % - The proportion of words pronounced correctly out of total words.
  • Whisper jiwer vs Whisper LLM - Two judging methodologies. jiwer uses raw Whisper-decoded transcripts; LLM-judged uses a follow-on LLM to normalize transcription noise. Both report the same metric family; LLM-judged tends to give lower error rates by reducing false positives from punctuation/casing.

For Pronunciation and WER, the residual gap on Lightning v3.1 (Standard) is concentrated in proper-noun rendering. Use a pronunciation dictionary to pin names, brands, and acronyms; with the dictionary applied, both metrics close to parity.