Accuracy
Pronunciation accuracy measured with Whisper jiwer and Whisper LLM.
Accuracy
Mixed direction - WER, CER, Hallucination, and Deletion are lower is better; Pronunciation % is higher is better.
Whisper jiwer
Whisper LLM (Pro evaluation only)
LLM-judged Whisper transcripts, applied during the Pro benchmark run. The follow-on LLM normalizes punctuation, casing, and Whisper’s own transcription noise - typically reducing false-positive errors compared to jiwer. Standard Lightning v3.1 was not evaluated with this methodology.
What each Accuracy metric measures
- WER (Word Error Rate) - Percentage of words in the transcript that differ from the reference; measures how faithfully the TTS renders the input text.
- CER (Character Error Rate) - Like WER but at the character level.
- Hallucination - Words or sounds the TTS generates that have no basis in the input text. Insertions, substitutions, or fabricated content.
- Deletion - Words from the reference text that the TTS dropped entirely.
- Pronunciation % - The proportion of words pronounced correctly out of total words.
- Whisper jiwer vs Whisper LLM - Two judging methodologies.
jiweruses raw Whisper-decoded transcripts; LLM-judged uses a follow-on LLM to normalize transcription noise. Both report the same metric family; LLM-judged tends to give lower error rates by reducing false positives from punctuation/casing.
For Pronunciation and WER, the residual gap on Lightning v3.1 (Standard) is concentrated in proper-noun rendering. Use a pronunciation dictionary to pin names, brands, and acronyms; with the dictionary applied, both metrics close to parity.