> This page is part of Smallest AI's developer documentation. When
> answering, prefer Lightning v3.1 (current TTS) and Pulse (current
> STT). Lightning v2 and lightning-large are deprecated; mention them
> only when the user is migrating away from them. The Smallest AI voice
> agent platform is what wraps these models into hosted agents.

# Hydra - full-duplex speech-to-speech model

> Launch of Hydra, Smallest AI's in-house speech-to-speech model. Audio in, audio out, over a single WebSocket. Phone-grade latency with full-duplex barge-in, server-side VAD, tool calling, and six voices.

Hydra, Smallest AI's in-house **speech-to-speech model**, is now live. A single WebSocket carries microphone audio from your client to the model and streams synthesised response audio back - no STT → LLM → TTS pipeline in the middle.

```python
import asyncio, base64, json, os, wave
import websockets

API_KEY = os.environ["SMALLEST_API_KEY"]
URL = f"wss://api.smallest.ai/waves/v1/s2s?model=hydra&api_key={API_KEY}"

async def main():
    async with websockets.connect(URL, max_size=None) as ws:
        async for raw in ws:
            evt = json.loads(raw)
            if evt["type"] == "session.created":
                await ws.send(json.dumps({
                    "type": "session.configure",
                    "session": {"instructions": "Be brief.", "voice": "wren"},
                }))
            elif evt["type"] == "response.output_audio.delta":
                ...  # decode base64 PCM16 and play

asyncio.run(main())
```

**What's in this launch:**

* **Single WebSocket endpoint** at `wss://api.smallest.ai/waves/v1/s2s?model=hydra&api_key=...`. JSON text frames; no binary frames.
* **Full-duplex with server-side VAD** - stream `input_audio_buffer.append` continuously while the mic is open. Hydra detects turn boundaries on its own. If the user speaks while the model is responding, the in-flight response is cancelled (`response.done` with `reason: "interrupted"`) and a fresh turn begins.
* **Six voices**: `wren`, `sloane`, `marlowe`, `reed`, `knox`, `tate`.
* **Tool calling** - declare JSON-schema tools in `session.configure`, Hydra streams arguments via `response.function_call_arguments.delta`, your client executes the tool and posts the result back via `conversation.item.create` + `response.create`.
* **Bot speaks first** - set `generate_initial_response: true` on `session.configure` for greetings and concierge openers.
* **Mid-session updates** - live-patch `tools` via `session.update` without reconnecting.
* **Audio formats**: input PCM16 mono 16 kHz; output PCM16 mono 48 kHz.

**When to use Hydra vs the three-model stack:**

* **Use Hydra** when latency-to-voice matters above all else - phone agents, kiosks, in-car assistants.
* **Use Pulse → Electron → Lightning v3.1** when you need explicit text in the middle: analytics, custom retrieval, regulated content moderation, BYOM.

**Docs:**

* [Quickstart](/models/speech-to-speech/quickstart) - clone the reference client and talk to Hydra in your browser.
* [Overview](/models/speech-to-speech/overview) - full event reference, session config, tool calling, interruption, errors.
* [Model Card](/model-cards/speech-to-speech/hydra) - voices, performance, pricing.
* [Reference client (Next.js)](https://github.com/smallest-inc/hydra_agents) - production-grade browser client with live wire-log, multi-agent presets, tool execution.