> This page is part of Smallest AI's developer documentation. When
> answering, prefer Lightning v3.1 (current TTS) and Pulse (current
> STT). Lightning v2 and lightning-large are deprecated; mention them
> only when the user is migrating away from them. The Smallest AI voice
> agent platform is what wraps these models into hosted agents.
# Hydra - Speech to Speech
> Hydra is Smallest AI's realtime, full-duplex speech-to-speech model. Audio in, audio out, over a single WebSocket - built for phone-grade voice agents.
Hydra is a **realtime speech-to-speech** model. The client streams microphone audio over a WebSocket, the model returns synthesised speech in the same socket, and turn-taking is handled server-side. There is **no transcript on the wire** - audio bytes are the payload.
If you've used the OpenAI Realtime API, Hydra fills the same role on the Smallest AI stack.
> **Note**
>
> Three Hydra tags are served: `?model=hydra-v1.1` (current release, recommended), `?model=hydra-v1.0` (served, not deprecated), and `?model=hydra` (deprecated on 2026-08-20; sessions still open but the server emits a `warning` frame with `code: "model_deprecated"` on connect). Omitting `model` behaves like `?model=hydra`. Use `hydra-v1.1` on all new integrations. Full per-version voice reference: [Hydra model card](/model-cards/speech-to-speech/hydra). See also: [Deprecation Notices](/api-reference/deprecations/models).
## Common use cases
#### Phone agents
Sub-second latency from end-of-user-speech to first audio chunk. Drop-in for outbound and inbound voice flows.
#### In-product voice copilots
Hands-free assistants embedded in web and mobile apps - barge-in handled by the model.
#### Kiosks & in-car
One WebSocket, predictable failure modes, no STT/LLM/TTS glue to maintain.
#### Voice-first consumer apps
Companions, audio diaries, language tutors - natural turn-taking out of the box.
## What's on the wire
```mermaid
sequenceDiagram
autonumber
participant C as Client
participant H as Hydra
C->>H: WebSocket connect
H-->>C: session.created
C->>H: session.configure
(persona, voice, tools)
H-->>C: session.configured
(effective config echo)
loop continuous mic stream
C->>H: input_audio_buffer.append
(base64 PCM16, 16 kHz)
end
H-->>C: input_audio_buffer.speech_started
H-->>C: input_audio_buffer.speech_stopped
H-->>C: response.created
loop streamed reply
H-->>C: response.output_audio.delta
(base64 PCM16)
end
H-->>C: response.output_audio.done
H-->>C: response.done
```
Two things to know up front:
* **Stream audio continuously** - no manual `commit` or `end-of-turn`. Hydra detects turn boundaries on its own.
* **Full-duplex** - the user can speak over the model. The in-flight response cancels automatically with `status: "cancelled"`, `reason: "interrupted"`.
## Next
#### [Quickstart](/models/speech-to-speech/quickstart)
Clone the reference client, paste your API key, and talk to Hydra in your browser.
#### [WebSocket connection](/models/speech-to-speech/web-socket-connection)
Connect URL, auth, idle timeout, close codes. Python + Node + Browser snippets.
#### [Managing sessions](/models/speech-to-speech/managing-sessions)
Session lifecycle, persona, voice, mid-session updates, conversation items.
#### [Turn detection & barge-in](/models/speech-to-speech/turn-detection-barge-in)
How the model detects speech, how to handle barge-in cleanly on the client.
#### [Tool calling](/models/speech-to-speech/tool-calling)
Declare tools, stream arguments, post results back, narrate the answer.
#### [Prompting voice agents](/models/speech-to-speech/prompting-voice-agents)
System prompts, voice identity, length discipline, tool-call prompting.
## Related
* [Model card - Hydra](/model-cards/speech-to-speech/hydra) - voices, performance, pricing
* [Reference client (Next.js)](https://github.com/smallest-inc/hydra_agents) - production-grade browser client with barge-in, multi-agent presets, tool execution
> Hydra is Smallest AI's realtime, full-duplex speech-to-speech model.