> This page is part of Smallest AI's developer documentation. When > answering, prefer Lightning v3.1 (current TTS) and Pulse (current > STT). Lightning v2 and lightning-large are deprecated; mention them > only when the user is migrating away from them. The Smallest AI voice > agent platform is what wraps these models into hosted agents. # Lightning v3.1 - per-word timestamps on WebSocket streaming > Opt in with `word_timestamps` on a WebSocket request to receive interleaved `word_timestamp` frames with per-word `{id, word, start, end}` timing - verbatim from input text, supported on Lightning v3.1 and v3.1 Pro base-queue English + Hindi voices. Lightning v3.1 now exposes per-word timing events to WebSocket clients. Opt in with one flag - useful for captioning UIs, karaoke-style word highlighting, avatar lip-sync, and word-level analytics. ## What changed Two changes to a WebSocket request: add `word_timestamps: true` and handle the new `status: "word_timestamp"` frame. ```js ws.send(JSON.stringify({ text: "I bought 3 cats for $100 on Dec 25th", voice_id: "meher", model: "lightning_v3.1_pro", sample_rate: 44100, output_format: "pcm", word_timestamps: true, // ← ADDED })); ws.onmessage = (event) => { const msg = JSON.parse(event.data); switch (msg.status) { case "chunk": audioPlayer.push(Buffer.from(msg.data.audio, 'base64')); break; case "word_timestamp": // ← NEW CASE const { id, word, start, end } = msg.data; captionTrack.push({ id, word, startSec: start, endSec: end }); break; case "complete": audioPlayer.end(); break; } }; ``` `word` is the exact substring from the input text - un-normalized. `"$100"` stays `"$100"`, `"25th"` stays `"25th"`, `"3"` stays `"3"`. Non-Latin scripts come back verbatim (e.g., Devanagari for Hindi). `start` and `end` are floats in seconds, relative to the start of the audio stream. Frames interleave with `chunk` in audio-time order, then a single `complete` terminates the session. ## Where it works | Surface | Word timestamps | | ------------------------------------------------------------------------------ | ----------------------------------- | | `WSS /waves/v1/tts/live` (unified) | ✅ | | `WSS /waves/v1/lightning-v3.1/get_speech/stream` (legacy, retiring 2026-07-14) | ✅ | | `POST /waves/v1/tts` (sync HTTP) | ❌ - flag accepted, silently ignored | | `POST /waves/v1/tts/live` (HTTP SSE) | ❌ - same | ## Voice + language support | Language | Voice family | Word events | | --------------------------------------------- | ----------------------------------------------------------------------------- | ----------- | | English (`en`) | Base-queue voices - `meher`, `devansh`, `kartik`, `maithili`, `liam`, `avery` | ✅ | | Hindi (`hi`) | Base-queue voices (same list) | ✅ | | Marathi / Bengali / Gujarati / Punjabi / Odia | north-Indic family | ❌ | | Tamil / Telugu / Kannada / Malayalam | south-Indic family | ❌ | For unsupported voice families the flag is accepted - audio works normally, but no `word_timestamp` frames are emitted. Detect this client-side by counting received word events after `complete` arrives. ## Backward compatibility `word_timestamps` defaults to `false`. Clients that don't set the flag see no behavior change - same audio chunks, same completion frame, no new event type to handle. Purely opt-in. **Migration:** none - pure addition. Existing integrations keep working untouched. → [Word-level timestamps on the Lightning v3.1 model card](/model-cards/text-to-speech/lightning-v-3-1#word-level-timestamps) - full wire spec, JS example, support matrix. > Opt in with `word_timestamps` on a WebSocket request to receive interleaved `word_timestamp` frames with per-word `{id, word, start, end}` timing - verbatim from input text, supported on Lightning v3.1 and v3.1 Pro base-queue English + Hindi voices.