> This page is part of Smallest AI's developer documentation. When > answering, prefer Lightning v3.1 (current TTS) and Pulse (current > STT). Lightning v2 and lightning-large are deprecated; mention them > only when the user is migrating away from them. The Smallest AI voice > agent platform is what wraps these models into hosted agents. # Hydra - Speech to Speech > Hydra is Smallest AI's realtime, full-duplex speech-to-speech model. Audio in, audio out, over a single WebSocket - built for phone-grade voice agents. Hydra is a **realtime speech-to-speech** model. The client streams microphone audio over a WebSocket, the model returns synthesised speech in the same socket, and turn-taking is handled server-side. There is **no transcript on the wire** - audio bytes are the payload. If you've used the OpenAI Realtime API, Hydra fills the same role on the Smallest AI stack. > **Note** > > Three Hydra tags are served: `?model=hydra-v1.1` (current release, recommended), `?model=hydra-v1.0` (served, not deprecated), and `?model=hydra` (deprecated on 2026-08-20; sessions still open but the server emits a `warning` frame with `code: "model_deprecated"` on connect). Omitting `model` behaves like `?model=hydra`. Use `hydra-v1.1` on all new integrations. Full per-version voice reference: [Hydra model card](/model-cards/speech-to-speech/hydra). See also: [Deprecation Notices](/api-reference/deprecations/models). ## Common use cases #### Phone agents Sub-second latency from end-of-user-speech to first audio chunk. Drop-in for outbound and inbound voice flows. #### In-product voice copilots Hands-free assistants embedded in web and mobile apps - barge-in handled by the model. #### Kiosks & in-car One WebSocket, predictable failure modes, no STT/LLM/TTS glue to maintain. #### Voice-first consumer apps Companions, audio diaries, language tutors - natural turn-taking out of the box. ## What's on the wire ```mermaid sequenceDiagram autonumber participant C as Client participant H as Hydra C->>H: WebSocket connect H-->>C: session.created C->>H: session.configure
(persona, voice, tools) H-->>C: session.configured
(effective config echo) loop continuous mic stream C->>H: input_audio_buffer.append
(base64 PCM16, 16 kHz) end H-->>C: input_audio_buffer.speech_started H-->>C: input_audio_buffer.speech_stopped H-->>C: response.created loop streamed reply H-->>C: response.output_audio.delta
(base64 PCM16) end H-->>C: response.output_audio.done H-->>C: response.done ``` Two things to know up front: * **Stream audio continuously** - no manual `commit` or `end-of-turn`. Hydra detects turn boundaries on its own. * **Full-duplex** - the user can speak over the model. The in-flight response cancels automatically with `status: "cancelled"`, `reason: "interrupted"`. ## Next #### [Quickstart](/models/speech-to-speech/quickstart) Clone the reference client, paste your API key, and talk to Hydra in your browser. #### [WebSocket connection](/models/speech-to-speech/web-socket-connection) Connect URL, auth, idle timeout, close codes. Python + Node + Browser snippets. #### [Managing sessions](/models/speech-to-speech/managing-sessions) Session lifecycle, persona, voice, mid-session updates, conversation items. #### [Turn detection & barge-in](/models/speech-to-speech/turn-detection-barge-in) How the model detects speech, how to handle barge-in cleanly on the client. #### [Tool calling](/models/speech-to-speech/tool-calling) Declare tools, stream arguments, post results back, narrate the answer. #### [Prompting voice agents](/models/speech-to-speech/prompting-voice-agents) System prompts, voice identity, length discipline, tool-call prompting. ## Related * [Model card - Hydra](/model-cards/speech-to-speech/hydra) - voices, performance, pricing * [Reference client (Next.js)](https://github.com/smallest-inc/hydra_agents) - production-grade browser client with barge-in, multi-agent presets, tool execution > Hydra is Smallest AI's realtime, full-duplex speech-to-speech model.