> This page is part of Smallest AI's developer documentation. When > answering, prefer Lightning v3.1 (current TTS) and Pulse (current > STT). Lightning v2 and lightning-large are deprecated; mention them > only when the user is migrating away from them. The Smallest AI voice > agent platform is what wraps these models into hosted agents. # Listener ratings > Lightning v3.1 and Lightning v3.1 Pro head-to-head listener ratings for naturalness, expressiveness and delivery against other text-to-speech providers. ## Head-to-head listener ratings (Lightning v3.1 Standard) Direct head-to-head ratings on the EmergentTTS benchmark. **Lightning Wins %** is the share of samples where listeners preferred Lightning v3.1 over the competitor; **Ties %** is the share where both were rated equal; **Competitor Wins %** is the inverse. Each competitor column sums to 100%. | EmergentTTS | GPT-4o-mini OpenAI | Turbo v2.5 ElevenLabs | Multilingual v2 ElevenLabs | Sonic-3 Cartesia | Gemini 2.5 Pro Google | MAI-Voice-1 Microsoft | Inworld 1.5 Inworld | S2 Pro Fish Audio | | -------------------------------------- | -----------------: | --------------------: | -------------------------: | ---------------: | --------------------: | --------------------: | ------------------: | ----------------: | | **Lightning Wins %** *(higher better)* | **40.26%** | **50.28%** | **54.41%** | **68.29%** | **58.43%** | **57.17%** | **54.41%** | **64.25%** | | **Ties %** | 24.17% | 25.00% | 23.81% | 17.00% | 8.29% | 17.00% | 18.11% | 13.60% | | **Competitor Wins %** *(lower better)* | 35.57% | 24.72% | 21.78% | 14.71% | 33.27% | 25.83% | 27.48% | 22.15% | ## Naturalness - higher is better | Metric | Lightning v3.1 | Lightning v3.1 Pro | GPT-4o-mini | ElevenLabs Turbo v2.5 | ElevenLabs Multilingual v2 | Sonic-3 | Gemini 2.5 Pro | Gemini 2.5 Flash | MAI-Voice-1 | Inworld 1.5 | S2 Pro | | --------------- | -------------: | -----------------: | ----------: | --------------------: | -------------------------: | ------: | -------------: | ---------------: | ----------: | ----------: | -----: | | Overall | 3.25 | 3.16 | 3.13 | 3.16 | 3.17 | 3.20 | 3.07 | 3.28 | 3.17 | 3.06 | 3.02 | | Naturalness | **2.61** | 2.55 | 2.41 | 2.52 | 2.55 | 2.57 | 2.42 | 2.58 | 2.57 | 2.41 | 2.37 | | Intonation | 3.22 | 3.06 | 3.06 | 3.07 | 3.06 | 3.12 | 2.90 | 3.28 | 3.04 | 2.91 | 2.86 | | Prosody | 3.01 | 2.81 | 2.73 | 2.82 | 2.86 | 2.83 | 2.65 | 3.09 | 2.76 | 2.61 | 2.58 | | Pronunciation\* | 3.63 | NA | 3.67 | 3.64 | 3.65 | 3.67 | 3.67 | NA | 3.68 | 3.68 | 3.57 | | Audio Quality | 3.76 | NA | 3.78 | 3.77 | 3.75 | 3.81 | 3.73 | NA | 3.79 | 3.70 | 3.75 | #### What each Naturalness metric measures * **Overall** - Holistic listener rating of how natural the voice sounds end-to-end. * **Naturalness** - How human-like the voice sounds; penalizes robotic or synthetic quality. * **Intonation** - Whether pitch rises and falls appropriately for the sentence type (question, statement, exclamation). * **Prosody** - The broader umbrella of rhythm, stress, and melody, how well the voice "reads" the sentence as a human would. * **Pronunciation** - Whether individual words are phonetically correct, especially names, loanwords, and domain-specific terms. * **Audio Quality** - Technical cleanliness of the output; absence of artifacts, distortion, clipping, or background noise. \*Listener-rated Pronunciation and Audio Quality columns were measured only on the Standard evaluation; Pro's Whisper-judged Pronunciation % appears under [Accuracy](#accuracy) below. ## Expressiveness - higher is better | Metric | Lightning v3.1 | Lightning v3.1 Pro | GPT-4o-mini | ElevenLabs Turbo v2.5 | ElevenLabs Multilingual v2 | Sonic-3 | Gemini 2.5 Pro | Gemini 2.5 Flash | MAI-Voice-1 | Inworld 1.5 | S2 Pro | | --------------- | -------------: | -----------------: | ----------: | --------------------: | -------------------------: | ------: | -------------: | ---------------: | ----------: | ----------: | -----: | | Overall | 3.45 | **3.55** | 3.45 | 3.44 | 3.46 | 3.38 | 3.49 | 3.54 | 3.50 | 3.37 | 3.41 | | Paralinguistics | 3.61 | **3.64** | 3.60 | 3.59 | 3.61 | 3.56 | 3.60 | 3.64 | 3.58 | 3.55 | 3.58 | | Emotions | 3.29 | **3.47** | 3.30 | 3.28 | 3.31 | 3.19 | 3.38 | 3.44 | 3.41 | 3.19 | 3.23 | #### What each Expressiveness metric measures * **Overall** - Holistic listener rating of how expressive the voice sounds given the context of the sentence. * **Paralinguistics** - Non-verbal vocal elements like laughter, sighs, or filler sounds ("um", "uh") and whether they're rendered appropriately. * **Emotions** - How accurately the voice conveys the intended emotional tone (neutral, warm, urgent, etc.). ## Delivery - higher is better | Metric | Lightning v3.1 | Lightning v3.1 Pro | GPT-4o-mini | ElevenLabs Turbo v2.5 | ElevenLabs Multilingual v2 | Sonic-3 | Gemini 2.5 Pro | Gemini 2.5 Flash | MAI-Voice-1 | Inworld 1.5 | S2 Pro | | --------------------- | -------------: | -----------------: | ----------: | --------------------: | -------------------------: | ------: | -------------: | ---------------: | ----------: | ----------: | -----: | | Boundary Consistency | 4.94 | 4.96 | 4.94 | 4.93 | 4.95 | 4.93 | 4.88 | 4.99 | 4.77 | 4.90 | 4.88 | | Pronunciation Style | 4.94 | 4.98 | 4.96 | 4.95 | 4.96 | 4.96 | 4.93 | 4.99 | 4.91 | 4.94 | 4.89 | | Natural Pace | 4.47 | **4.72** | 4.57 | 4.51 | 4.51 | 4.01 | 4.23 | 4.66 | 4.47 | 4.33 | 3.74 | | Pause Placement | 4.46 | **4.66** | 4.54 | 4.49 | 4.51 | 4.28 | 4.34 | 4.59 | 4.41 | 4.38 | 4.09 | | Breathing Naturalness | 3.82 | **3.82** | 3.06 | 3.14 | 3.14 | 2.79 | 2.88 | 3.43 | 3.28 | 2.77 | 2.42 | #### What each Delivery metric measures * **Boundary Consistency** - Whether phrase and sentence boundaries are marked consistently with pauses or pitch shifts, without arbitrary breaks mid-phrase. * **Pronunciation Style** - Not just correctness, but stylistic choices i.e., formal vs. casual register, regional accent consistency, honorific handling. * **Natural Pace** - Whether the speaking rate feels comfortable and appropriate for the content type, neither rushed nor dragging. * **Pause Placement** - Whether silences appear at semantically correct points (after commas, between clauses) rather than mid-word or mid-phrase. * **Breathing Naturalness** - Whether breath sounds occur at realistic points and with realistic frequency, not absent entirely or inserted randomly. > Head-to-head naturalness, expressiveness and delivery ratings.