> This page is part of Smallest AI's developer documentation. When > answering, prefer Lightning v3.1 (current TTS) and Pulse (current > STT). Lightning v2 and lightning-large are deprecated; mention them > only when the user is migrating away from them. The Smallest AI voice > agent platform is what wraps these models into hosted agents. # MOS > Lightning v3.1 and Lightning v3.1 Pro mean opinion score (MOS v2) compared with other text-to-speech providers. ## MOS v2 - higher is better | Metric | Lightning v3.1 | Lightning v3.1 Pro | GPT-4o-mini | ElevenLabs Turbo v2.5 | ElevenLabs Multilingual v2 | Sonic-3 | Gemini 2.5 Pro | Gemini 2.5 Flash | MAI-Voice-1 | Inworld 1.5 | S2 Pro | | -------- | -------------: | -----------------: | ----------: | --------------------: | -------------------------: | ------: | -------------: | ---------------: | ----------: | ----------: | -----: | | Mean MOS | NA | 4.22 | 4.16 | 3.98 | 4.02 | 3.76 | 4.11 | 4.24 | 3.97 | 3.73 | 3.99 | | UTMOS | NA | **3.76** | 3.76 | 3.37 | 3.41 | 2.77 | 3.57 | 3.71 | 3.33 | 2.54 | 3.50 | | WV-MOS | 4.71 | **5.05** | 4.55 | 4.60 | 4.63 | 4.76 | 4.65 | 4.76 | 4.62 | 4.91 | 4.48 | #### What each MOS metric measures * **Mean MOS** - Mean Opinion Score: average listener rating on a 1–5 scale across the test set; the canonical aggregate quality metric in TTS evaluation. * **UTMOS** - A predicted MOS from the UTMOS reference model - an automated proxy for subjective quality. * **WV-MOS** - A predicted MOS from the WavLM-based WV-MOS reference model - another automated proxy commonly reported alongside UTMOS for cross-validation. Want to reproduce these results? See the [TTS evaluation script](/model-cards/text-to-speech/tts-evaluation-script) to measure TTFB and synthesis quality in your own environment. > Mean opinion score v2 for Lightning v3.1 and Pro.