> This page is part of Smallest AI's developer documentation. When > answering, prefer Lightning v3.1 (current TTS) and Pulse (current > STT). Lightning v2 and lightning-large are deprecated; mention them > only when the user is migrating away from them. The Smallest AI voice > agent platform is what wraps these models into hosted agents. # Models ## Entries - [Short-lived access tokens for TTS, STT and speech-to-speech](https://docs.smallest.ai/changelog/models/2026/9/22.md): Mint a `wat_` token on your server with `POST /waves/v1/auth/token` and let browser or mobile clients call TTS, STT and speech-to-speech with it. The API key never leaves your backend. - [WebSocket keep-alive frame is now documented](https://docs.smallest.ai/changelog/models/2026/9/10.md): Send `{"type":"ping"}` on any Waves WebSocket to hold an idle connection open and get a `pong` back. Available since August; the docs and API reference now cover it, along with the real session limits. - [Docs: fix five stale internal links after slug renames](https://docs.smallest.ai/changelog/models/2026/9/7.md): Five internal links in the Models docs now point at the correct pages: - Electron model card API-reference link (electron-chat-completions) - LLM overview + q - [Speech-to-Speech: Hydra V1.1 released, `?model=hydra` deprecated](https://docs.smallest.ai/changelog/models/2026/8/20.md): Hydra V1.1 is the current release of the realtime speech-to-speech model. - [STT: new language `hi-dev` (Devanagari-only Hindi)](https://docs.smallest.ai/changelog/models/2026/8/16.md): Pulse STT streaming supports a new Hindi language code, hi-dev, that transcribes Hindi audio using a pure-Hindi model and writes everything in Devanagari. - [Speech to Text: VAD events on the live WebSocket API reference](https://docs.smallest.ai/changelog/models/2026/8/10.md): The Pulse STT live WebSocket now models the acoustic voice-activity events in the API reference: speech_started and speech_ended, emitted alongside. - [Waves: analytics endpoints on the API reference](https://docs.smallest.ai/changelog/models/2026/8/8.md): Nine previously undocumented Waves analytics endpoints now render on the API reference under a new Analytics section: - [STT API-ref: response examples now match the real API](https://docs.smallest.ai/changelog/models/2026/8/6.md): Follow-up to yesterday's word_timestamps clarification. - [STT API: clarify that words[] and utterances[] need word_timestamps=true](https://docs.smallest.ai/changelog/models/2026/8/5.md): The STT API-ref example response showed populated words[] and utterances[] arrays, but a default request from the docs explorer or Postman returns both. - [New `x-expire-content` header: opt a request into content deletion (enterprise)](https://docs.smallest.ai/changelog/models/2026/8/4.md): Documented as a header parameter on the HTTP routes - POST /waves/v1/tts, POST /waves/v1/tts/live, POST /waves/v1/stt/ - and on the WebSocket routes. - [TTS: latency measurement caveat + legacy /streaming-tts/stream artifacts retired](https://docs.smallest.ai/changelog/models/2026/8/3.md): Two related updates to the TTS docs. - [TTS language surface aligned with product-team canonical](https://docs.smallest.ai/changelog/models/2026/7/21.md): The unified TTS endpoints (POST /waves/v1/tts + wss://api.smallest.ai/waves/v1/tts/live) now document the full language surface the platform accepts. - [Models docs URL migration - waves → models, version dropdown flattened](https://docs.smallest.ai/changelog/models/2026/7/18.md): The Models product now lives under /models/ on the docs site, replacing the older /waves/ prefix. - [Waves docs v2.2.0 and v3.0.1 retired; v4.0.0 is now the sole documented surface](https://docs.smallest.ai/changelog/models/2026/7/17.md): The v4.0.0 doc surface is now the only supported view of the Waves models. - [Agno integration guide](https://docs.smallest.ai/changelog/models/2026/7/16.md): Added a documentation page for generating speech with Lightning v3.1 and Lightning v3.1 Pro inside Agno, the open-source Python framework for building. - [Pulse STT: `vad_events` query parameter on the streaming WebSocket](https://docs.smallest.ai/changelog/models/2026/6/23.md): The Pulse STT WebSocket accepts a new query parameter, vad_events. - [Lightning v3.1 + Pulse STT API-ref Python snippets - `SmallestAI(api_key=...)` (was `token=...`)](https://docs.smallest.ai/changelog/models/2026/6/12.md): Doc-only update - the Python snippets in the Lightning v3.1 and Pulse STT API references were still using the pre-4.4.5 `SmallestAI(token=...)` kwarg. Replaced with `api_key=...` to match the current SDK. - [WebSocket auth and default URL fixes](https://docs.smallest.ai/changelog/models/2026/6/3.md): Fixed WebSocket auth and URLs; deprecated assorted `v2.2.0` and `v3.0.1` endpoints. - [Hydra - full-duplex speech-to-speech model](https://docs.smallest.ai/changelog/models/2026/5/23.md): Launch of Hydra, Smallest AI's in-house speech-to-speech model. Audio in, audio out, over a single WebSocket. Phone-grade latency with full-duplex barge-in, server-side VAD, tool calling, and six voices. - [Electron LLM - chat completions on the Waves API](https://docs.smallest.ai/changelog/models/2026/5/22.md): Launch of Electron, Smallest AI's in-house language model. OpenAI-compatible chat completions, sub-300 ms TTFT, 70 languages with first-class Indic support, voice-agent-optimized tool calling, and automatic prefix caching. - [API reference rewrite - Lightning v3.1 + Pulse](https://docs.smallest.ai/changelog/models/2026/5/12.md): Lightning v3.1 (sync / SSE / WS) and Pulse (REST / WS) endpoint descriptions rewritten with copy-paste-runnable examples in cURL, Python (smallestai>=4.4.0), and JavaScript (fetch / ws). All samples live-tested against api.smallest.ai. - [Waves API spec - corrected `output_format` enum and resynced base ↔ v4 overrides](https://docs.smallest.ai/changelog/models/2026/5/7.md): The output_format enum on Lightning v3.1 (POST /waves/v1/lightning-v3.1/get_speech and /stream) is now correctly documented as ['pcm', 'mp3', 'wav'. - [OpenWhispr integration guide](https://docs.smallest.ai/changelog/models/2026/4/22.md): Added a documentation page for using Smallest AI's Pulse model as the speech-to-text engine inside OpenWhispr, the open-source desktop dictation app for. - [Unified voice cloning API](https://docs.smallest.ai/changelog/models/2026/4/20.md): Voice cloning now has a single cross-model endpoint: POST /waves/v1/voice-cloning.