Skip to navigation

Listener ratings

Head-to-head naturalness, expressiveness and delivery ratings.
View as Markdown

Head-to-head listener ratings (Lightning v3.1 Standard)

Direct head-to-head ratings on the EmergentTTS benchmark. Lightning Wins % is the share of samples where listeners preferred Lightning v3.1 over the competitor; Ties % is the share where both were rated equal; Competitor Wins % is the inverse. Each competitor column sums to 100%.

EmergentTTSGPT-4o-mini
OpenAI
Turbo v2.5
ElevenLabs
Multilingual v2
ElevenLabs
Sonic-3
Cartesia
Gemini 2.5 Pro
Google
MAI-Voice-1
Microsoft
Inworld 1.5
Inworld
S2 Pro
Fish Audio
Lightning Wins % (higher better)40.26%50.28%54.41%68.29%58.43%57.17%54.41%64.25%
Ties %24.17%25.00%23.81%17.00%8.29%17.00%18.11%13.60%
Competitor Wins % (lower better)35.57%24.72%21.78%14.71%33.27%25.83%27.48%22.15%

Naturalness - higher is better

MetricLightning v3.1Lightning v3.1 ProGPT-4o-miniElevenLabs Turbo v2.5ElevenLabs Multilingual v2Sonic-3Gemini 2.5 ProGemini 2.5 FlashMAI-Voice-1Inworld 1.5S2 Pro
Overall3.253.163.133.163.173.203.073.283.173.063.02
Naturalness2.612.552.412.522.552.572.422.582.572.412.37
Intonation3.223.063.063.073.063.122.903.283.042.912.86
Prosody3.012.812.732.822.862.832.653.092.762.612.58
Pronunciation*3.63NA3.673.643.653.673.67NA3.683.683.57
Audio Quality3.76NA3.783.773.753.813.73NA3.793.703.75
  • Overall - Holistic listener rating of how natural the voice sounds end-to-end.
  • Naturalness - How human-like the voice sounds; penalizes robotic or synthetic quality.
  • Intonation - Whether pitch rises and falls appropriately for the sentence type (question, statement, exclamation).
  • Prosody - The broader umbrella of rhythm, stress, and melody, how well the voice “reads” the sentence as a human would.
  • Pronunciation - Whether individual words are phonetically correct, especially names, loanwords, and domain-specific terms.
  • Audio Quality - Technical cleanliness of the output; absence of artifacts, distortion, clipping, or background noise.
*Listener-rated Pronunciation and Audio Quality columns were measured only on the Standard evaluation; Pro’s Whisper-judged Pronunciation % appears under Accuracy below.

Expressiveness - higher is better

MetricLightning v3.1Lightning v3.1 ProGPT-4o-miniElevenLabs Turbo v2.5ElevenLabs Multilingual v2Sonic-3Gemini 2.5 ProGemini 2.5 FlashMAI-Voice-1Inworld 1.5S2 Pro
Overall3.453.553.453.443.463.383.493.543.503.373.41
Paralinguistics3.613.643.603.593.613.563.603.643.583.553.58
Emotions3.293.473.303.283.313.193.383.443.413.193.23
  • Overall - Holistic listener rating of how expressive the voice sounds given the context of the sentence.
  • Paralinguistics - Non-verbal vocal elements like laughter, sighs, or filler sounds (“um”, “uh”) and whether they’re rendered appropriately.
  • Emotions - How accurately the voice conveys the intended emotional tone (neutral, warm, urgent, etc.).

Delivery - higher is better

MetricLightning v3.1Lightning v3.1 ProGPT-4o-miniElevenLabs Turbo v2.5ElevenLabs Multilingual v2Sonic-3Gemini 2.5 ProGemini 2.5 FlashMAI-Voice-1Inworld 1.5S2 Pro
Boundary Consistency4.944.964.944.934.954.934.884.994.774.904.88
Pronunciation Style4.944.984.964.954.964.964.934.994.914.944.89
Natural Pace4.474.724.574.514.514.014.234.664.474.333.74
Pause Placement4.464.664.544.494.514.284.344.594.414.384.09
Breathing Naturalness3.823.823.063.143.142.792.883.433.282.772.42
  • Boundary Consistency - Whether phrase and sentence boundaries are marked consistently with pauses or pitch shifts, without arbitrary breaks mid-phrase.
  • Pronunciation Style - Not just correctness, but stylistic choices i.e., formal vs. casual register, regional accent consistency, honorific handling.
  • Natural Pace - Whether the speaking rate feels comfortable and appropriate for the content type, neither rushed nor dragging.
  • Pause Placement - Whether silences appear at semantically correct points (after commas, between clauses) rather than mid-word or mid-phrase.
  • Breathing Naturalness - Whether breath sounds occur at realistic points and with realistic frequency, not absent entirely or inserted randomly.