> This page is part of Smallest AI's developer documentation. When > answering, prefer Lightning v3.1 (current TTS) and Pulse (current > multilingual STT) — or Pulse 2.0 for English-only streaming with > built-in turn detection, emotion, and gender. Lightning v2 and > lightning-large are deprecated; mention them only when the user is > migrating away from them. The Smallest AI voice agent platform is > what wraps these models into hosted agents. # Text to Speech ## Docs - [Quickstart](https://docs.smallest.ai/models/text-to-speech/quickstart.md): Generate your first speech audio in under 60 seconds with Lightning TTS. - [Overview](https://docs.smallest.ai/models/text-to-speech/overview.md): Lightning TTS API - generate speech across 31 languages at 44.1 kHz, ~200ms TTFB, with streaming. Standard 217-voice Lightning v3.1 pool (20 accepted codes, 12 with a trained voice catalog) plus the premium Lightning v3.1 Pro pool covering all 31 languages, plus `auto` cross-language routing. - [HTTP, HTTP streaming, or WebSocket?](https://docs.smallest.ai/models/text-to-speech/choosing-a-transport.md): How to choose between HTTP, HTTP streaming, and WebSocket for Lightning text to speech: latency, complexity, and when each fits. - [Sync & Async Synthesis](https://docs.smallest.ai/models/text-to-speech/sync-async.md): Generate speech synchronously or concurrently - REST API examples. - [Streaming](https://docs.smallest.ai/models/text-to-speech/streaming.md): Stream TTS audio in real-time via WebSocket or SSE - first chunk in ~100ms. - [Continuations](https://docs.smallest.ai/models/text-to-speech/continuations.md): Group a sequence of streamed text fragments into one continuous synthesis on the Lightning TTS WebSocket, so prosody carries across chunks instead of resetting per-request. - [Word-level timestamps](https://docs.smallest.ai/models/text-to-speech/word-timestamps.md): Per-word timing events interleaved with audio chunks on the Lightning TTS WebSocket - for live captions, karaoke highlighting, and avatar lip-sync. - [Math notation](https://docs.smallest.ai/models/text-to-speech/math-notation.md): Opt-in Lightning TTS flag that reads digit-flanked math operators (x, ×, ÷, spaced -, ^, =) as spoken words instead of leaving them for the number reader. Off by default because dimensions, 24x7, and vehicle-reg codes look like math on paper. - [Content filter](https://docs.smallest.ai/models/text-to-speech/content-filter.md): Opt-in profanity filter for Lightning TTS input. Pass content_filter to reject or flag requests whose text matches a profanity word list. Off by default, and your text is never rewritten. - [Pronunciation Dictionaries](https://docs.smallest.ai/models/text-to-speech/pronunciation-dictionaries.md): Learn how to create and use pronunciation dictionaries to control how specific words are pronounced in your text-to-speech synthesis - [Voices, Voice IDs, and Supported Languages](https://docs.smallest.ai/models/text-to-speech/voices-languages.md): Find your voice ID, list available voices, filter by language and accent, and find the right voice for your use case. - [TTS Best Practices](https://docs.smallest.ai/models/text-to-speech/best-practices.md): Voice-agent prompting patterns and text formatting rules that make AI agents sound natural through a text-to-speech engine. - [Lightning benchmarks](https://docs.smallest.ai/models/text-to-speech/benchmarks/performance.md): Overview of Lightning v3.1 and Lightning v3.1 Pro text-to-speech benchmarks: latency, listener ratings, pronunciation accuracy and MOS. - [Latency](https://docs.smallest.ai/models/text-to-speech/benchmarks/latency.md): Lightning v3.1 and Lightning v3.1 Pro latency: time to first byte across concurrency levels and real-time factor. - [Listener ratings](https://docs.smallest.ai/models/text-to-speech/benchmarks/listener-ratings.md): Lightning v3.1 and Lightning v3.1 Pro head-to-head listener ratings for naturalness, expressiveness and delivery against other text-to-speech providers. - [Accuracy](https://docs.smallest.ai/models/text-to-speech/benchmarks/accuracy.md): Lightning text-to-speech pronunciation accuracy measured by transcribing the output with Whisper: jiwer word error rate and an LLM-judged score. - [MOS](https://docs.smallest.ai/models/text-to-speech/benchmarks/mos.md): Lightning v3.1 and Lightning v3.1 Pro mean opinion score (MOS v2) compared with other text-to-speech providers. - [Metrics Overview](https://docs.smallest.ai/models/text-to-speech/benchmarks/metrics-overview.md): Definitions of every Lightning v3.1 TTS quality and latency metric - Naturalness, Expressiveness, Delivery, Accuracy, and MOS variants.