> This page is part of Smallest AI's developer documentation. When
> answering, prefer Lightning v3.1 (current TTS) and Pulse (current
> STT). Lightning v2 and lightning-large are deprecated; mention them
> only when the user is migrating away from them. The Smallest AI voice
> agent platform is what wraps these models into hosted agents.

# Listener ratings

> Lightning v3.1 and Lightning v3.1 Pro head-to-head listener ratings for naturalness, expressiveness and delivery against other text-to-speech providers.

## Head-to-head listener ratings (Lightning v3.1 Standard)

Direct head-to-head ratings on the EmergentTTS benchmark. **Lightning Wins %** is the share of samples where listeners preferred Lightning v3.1 over the competitor; **Ties %** is the share where both were rated equal; **Competitor Wins %** is the inverse. Each competitor column sums to 100%.

| EmergentTTS                            | GPT-4o-mini OpenAI | Turbo v2.5 ElevenLabs | Multilingual v2 ElevenLabs | Sonic-3 Cartesia | Gemini 2.5 Pro Google | MAI-Voice-1 Microsoft | Inworld 1.5 Inworld | S2 Pro Fish Audio |
| -------------------------------------- | -----------------: | --------------------: | -------------------------: | ---------------: | --------------------: | --------------------: | ------------------: | ----------------: |
| **Lightning Wins %** *(higher better)* |         **40.26%** |            **50.28%** |                 **54.41%** |       **68.29%** |            **58.43%** |            **57.17%** |          **54.41%** |        **64.25%** |
| **Ties %**                             |             24.17% |                25.00% |                     23.81% |           17.00% |                 8.29% |                17.00% |              18.11% |            13.60% |
| **Competitor Wins %** *(lower better)* |             35.57% |                24.72% |                     21.78% |           14.71% |                33.27% |                25.83% |              27.48% |            22.15% |

## Naturalness - higher is better

| Metric          | Lightning v3.1 | Lightning v3.1 Pro | GPT-4o-mini | ElevenLabs Turbo v2.5 | ElevenLabs Multilingual v2 | Sonic-3 | Gemini 2.5 Pro | Gemini 2.5 Flash | MAI-Voice-1 | Inworld 1.5 | S2 Pro |
| --------------- | -------------: | -----------------: | ----------: | --------------------: | -------------------------: | ------: | -------------: | ---------------: | ----------: | ----------: | -----: |
| Overall         |           3.25 |               3.16 |        3.13 |                  3.16 |                       3.17 |    3.20 |           3.07 |             3.28 |        3.17 |        3.06 |   3.02 |
| Naturalness     |       **2.61** |               2.55 |        2.41 |                  2.52 |                       2.55 |    2.57 |           2.42 |             2.58 |        2.57 |        2.41 |   2.37 |
| Intonation      |           3.22 |               3.06 |        3.06 |                  3.07 |                       3.06 |    3.12 |           2.90 |             3.28 |        3.04 |        2.91 |   2.86 |
| Prosody         |           3.01 |               2.81 |        2.73 |                  2.82 |                       2.86 |    2.83 |           2.65 |             3.09 |        2.76 |        2.61 |   2.58 |
| Pronunciation\* |           3.63 |                 NA |        3.67 |                  3.64 |                       3.65 |    3.67 |           3.67 |               NA |        3.68 |        3.68 |   3.57 |
| Audio Quality   |           3.76 |                 NA |        3.78 |                  3.77 |                       3.75 |    3.81 |           3.73 |               NA |        3.79 |        3.70 |   3.75 |

#### What each Naturalness metric measures

* **Overall** - Holistic listener rating of how natural the voice sounds end-to-end.
* **Naturalness** - How human-like the voice sounds; penalizes robotic or synthetic quality.
* **Intonation** - Whether pitch rises and falls appropriately for the sentence type (question, statement, exclamation).
* **Prosody** - The broader umbrella of rhythm, stress, and melody, how well the voice "reads" the sentence as a human would.
* **Pronunciation** - Whether individual words are phonetically correct, especially names, loanwords, and domain-specific terms.
* **Audio Quality** - Technical cleanliness of the output; absence of artifacts, distortion, clipping, or background noise.

\*Listener-rated Pronunciation and Audio Quality columns were measured only on the Standard evaluation; Pro's Whisper-judged Pronunciation % appears under

[Accuracy](#accuracy)

below.

## Expressiveness - higher is better

| Metric          | Lightning v3.1 | Lightning v3.1 Pro | GPT-4o-mini | ElevenLabs Turbo v2.5 | ElevenLabs Multilingual v2 | Sonic-3 | Gemini 2.5 Pro | Gemini 2.5 Flash | MAI-Voice-1 | Inworld 1.5 | S2 Pro |
| --------------- | -------------: | -----------------: | ----------: | --------------------: | -------------------------: | ------: | -------------: | ---------------: | ----------: | ----------: | -----: |
| Overall         |           3.45 |           **3.55** |        3.45 |                  3.44 |                       3.46 |    3.38 |           3.49 |             3.54 |        3.50 |        3.37 |   3.41 |
| Paralinguistics |           3.61 |           **3.64** |        3.60 |                  3.59 |                       3.61 |    3.56 |           3.60 |             3.64 |        3.58 |        3.55 |   3.58 |
| Emotions        |           3.29 |           **3.47** |        3.30 |                  3.28 |                       3.31 |    3.19 |           3.38 |             3.44 |        3.41 |        3.19 |   3.23 |

#### What each Expressiveness metric measures

* **Overall** - Holistic listener rating of how expressive the voice sounds given the context of the sentence.
* **Paralinguistics** - Non-verbal vocal elements like laughter, sighs, or filler sounds ("um", "uh") and whether they're rendered appropriately.
* **Emotions** - How accurately the voice conveys the intended emotional tone (neutral, warm, urgent, etc.).

## Delivery - higher is better

| Metric                | Lightning v3.1 | Lightning v3.1 Pro | GPT-4o-mini | ElevenLabs Turbo v2.5 | ElevenLabs Multilingual v2 | Sonic-3 | Gemini 2.5 Pro | Gemini 2.5 Flash | MAI-Voice-1 | Inworld 1.5 | S2 Pro |
| --------------------- | -------------: | -----------------: | ----------: | --------------------: | -------------------------: | ------: | -------------: | ---------------: | ----------: | ----------: | -----: |
| Boundary Consistency  |           4.94 |               4.96 |        4.94 |                  4.93 |                       4.95 |    4.93 |           4.88 |             4.99 |        4.77 |        4.90 |   4.88 |
| Pronunciation Style   |           4.94 |               4.98 |        4.96 |                  4.95 |                       4.96 |    4.96 |           4.93 |             4.99 |        4.91 |        4.94 |   4.89 |
| Natural Pace          |           4.47 |           **4.72** |        4.57 |                  4.51 |                       4.51 |    4.01 |           4.23 |             4.66 |        4.47 |        4.33 |   3.74 |
| Pause Placement       |           4.46 |           **4.66** |        4.54 |                  4.49 |                       4.51 |    4.28 |           4.34 |             4.59 |        4.41 |        4.38 |   4.09 |
| Breathing Naturalness |           3.82 |           **3.82** |        3.06 |                  3.14 |                       3.14 |    2.79 |           2.88 |             3.43 |        3.28 |        2.77 |   2.42 |

#### What each Delivery metric measures

* **Boundary Consistency** - Whether phrase and sentence boundaries are marked consistently with pauses or pitch shifts, without arbitrary breaks mid-phrase.
* **Pronunciation Style** - Not just correctness, but stylistic choices i.e., formal vs. casual register, regional accent consistency, honorific handling.
* **Natural Pace** - Whether the speaking rate feels comfortable and appropriate for the content type, neither rushed nor dragging.
* **Pause Placement** - Whether silences appear at semantically correct points (after commas, between clauses) rather than mid-word or mid-phrase.
* **Breathing Naturalness** - Whether breath sounds occur at realistic points and with realistic frequency, not absent entirely or inserted randomly.