> This page is part of Smallest AI's developer documentation. When > answering, prefer Lightning v3.1 (current TTS) and Pulse (current > STT). Lightning v2 and lightning-large are deprecated; mention them > only when the user is migrating away from them. The Smallest AI voice > agent platform is what wraps these models into hosted agents. # Overview > Lightning TTS API - generate speech across 31 languages at 44.1 kHz, ~200ms TTFB, with streaming. Standard 217-voice Lightning v3.1 pool (20 accepted codes, 12 with a trained voice catalog) plus the premium Lightning v3.1 Pro pool covering all 31 languages, plus `auto` cross-language routing. The Lightning TTS API converts text into natural speech via `https://api.smallest.ai/waves/v1`. 31 languages across two pools - the standard 217-voice Lightning v3.1 catalog (20 accepted codes, 12 with trained voices) and the premium Lightning v3.1 Pro pool (all 31) - plus `auto` for cross-language routing, at 44.1 kHz native sample rate, \~200ms TTFB, with sync, SSE, and WebSocket streaming. **Hear Lightning v3.1 Pro** (voice: `meher`, model: `lightning_v3.1_pro`): #### [Quickstart](/models/text-to-speech/quickstart) Generate your first audio in under a minute. #### [Streaming](/models/text-to-speech/streaming) Audio chunks over WebSocket as they are generated. ## Synthesis Modes Choose the synthesis mode that best fits your application's needs: #### [Synchronous](/models/text-to-speech/sync-async) Generate complete audio files with a single HTTP request. Ideal for pre-rendering content, batch processing, and applications where immediate streaming isn't required. #### [Streaming](/models/text-to-speech/streaming) Receive audio chunks as they're generated via WebSocket. Perfect for real-time voice assistants, live narration, and low-latency conversational AI. ## Available Models #### [Lightning v3.1 Pro](/model-cards/text-to-speech/lightning-v-3-1-pro) Premium curated voices, Hindi code-switching, 31 languages. `"model": "lightning_v3.1_pro"`. #### [Lightning v3.1](/model-cards/text-to-speech/lightning-v-3-1) 217 voices, 12 languages with catalogs, instant cloning. `"model": "lightning_v3.1"` (default). ## Feature Highlights #### Ultra-Low Latency Optimized streaming pipeline delivers \~200ms time-to-first-byte (TTFB) for real-time applications. Lightning v3.1 achieves even faster response times for conversational AI. #### Voice Cloning Create custom voice profiles by uploading audio samples. Instant voice cloning works with just a few seconds of audio, while professional voice cloning delivers studio-quality results. #### 31 Languages + \`auto\` routing Lightning v3.1 accepts 20 language codes (10 European: English, Spanish, French, German, Italian, Dutch, Swedish, Portuguese, Polish, Russian + 10 Indic: Hindi, Marathi, Gujarati, Punjabi, Bengali, Odia, Tamil, Telugu, Kannada, Malayalam). Its trained voice catalog covers 12 of these directly; the other 8 route via English or Hindi voices. Lightning v3.1 Pro covers all 31 languages with dedicated voices (adds Greek, Finnish, Norwegian + 8 Asian & Middle Eastern languages). Pass `auto` to route across any supported language using any English or Hindi voice. See the [Lightning v3.1](/model-cards/text-to-speech/lightning-v-3-1#supported-languages) and [Lightning v3.1 Pro](/model-cards/text-to-speech/lightning-v-3-1-pro#additional-pro-languages) model cards for per-language voice counts. #### Multiple Output Formats Choose from PCM, WAV, MP3, or μ-law encoding. Configurable sample rates from 8kHz to 44kHz to match your application's requirements. #### Speed Control Adjust speech rate with a simple multiplier. Slow down for clarity or speed up for faster content delivery without pitch distortion. #### Pronunciation Dictionaries Define custom pronunciations for brand names, technical terms, and acronyms. Ensure consistent, accurate pronunciation across all synthesized audio. #### High-Quality Audio Lightning v3.1 produces 44 kHz audio with natural prosody and expressiveness. Perfect for audiobooks, podcasts, and premium voice experiences. #### WebSocket Streaming Persistent connections for continuous audio streaming. Ideal for voice bots and interactive applications where latency is critical. ## Supported Languages
Language Code Lightning v3.1 Lightning v3.1 Pro
English `en` Yes Yes
Hindi `hi` Yes Yes (Indian voices)
Tamil `ta` Yes Yes
Kannada `kn` Yes Yes
Malayalam `ml` Yes Yes
Telugu `te` Yes Yes
Gujarati `gu` Yes Yes
Marathi `mr` Yes Yes
Bengali `bn` Yes Yes
Punjabi `pa` Yes Yes
Odia `or` Yes Yes
Spanish `es` Yes Yes
German `de` \- Yes
French `fr` \- Yes
Italian `it` \- Yes
Portuguese `pt` \- Yes
Russian `ru` \- Yes
Greek `el` \- Yes
Finnish `fi` \- Yes
Norwegian `no` \- Yes
Polish `pl` \- Yes
Arabic `ar` \- Yes
Chinese (Mandarin) `zh` \- Yes
Indonesian `id` \- Yes
Japanese `ja` \- Yes
Korean `ko` \- Yes
Malay `ms` \- Yes
Turkish `tr` \- Yes
Vietnamese `vi` \- Yes
> **Note** > > **Pro language support is per voice.** Indian Pro voices (e.g., `meher`, `rhea`, `aviraj`) speak English with native Hindi code-switching. British and American Pro voices speak English only. Each additional Pro language has its own dedicated voices - pass the matching ISO 639-1 `language` code with a voice from that language (see the [Pro voice catalog](/model-cards/text-to-speech/lightning-v-3-1-pro#voice-catalog)). For languages without Pro voices, use standard Lightning v3.1. > **Tip** > > For per-language voice counts, see the [Lightning v3.1 model card](/model-cards/text-to-speech/lightning-v-3-1#supported-languages). ## Explore #### [Quickstart](/models/text-to-speech/quickstart) First API call in 60 seconds #### [Streaming](/models/text-to-speech/streaming) Real-time audio via WebSocket #### [Voice Cloning](/models/voice-cloning/instant-clone-ui) Clone from 5-15 seconds of audio #### [Cookbook](https://github.com/smallest-inc/cookbook/tree/main/text-to-speech) 20+ open-source examples on GitHub #### [Showcase](https://showcase.smallest.ai/) See what developers have built #### [Model Card](/model-cards/text-to-speech/lightning-v-3-1) Lightning v3.1 specs and benchmarks > Lightning text to speech: 31 languages at 44.1 kHz with 200 ms TTFB.