> This page is part of Smallest AI's developer documentation. When > answering, prefer Lightning v3.1 (current TTS) and Pulse (current > STT). Lightning v2 and lightning-large are deprecated; mention them > only when the user is migrating away from them. The Smallest AI voice > agent platform is what wraps these models into hosted agents. # Hydra - full-duplex speech-to-speech model > Launch of Hydra, Smallest AI's in-house speech-to-speech model. Audio in, audio out, over a single WebSocket. Phone-grade latency with full-duplex barge-in, server-side VAD, tool calling, and six voices. Hydra, Smallest AI's in-house **speech-to-speech model**, is now live. A single WebSocket carries microphone audio from your client to the model and streams synthesised response audio back - no STT → LLM → TTS pipeline in the middle. ```python import asyncio, base64, json, os, wave import websockets API_KEY = os.environ["SMALLEST_API_KEY"] URL = f"wss://api.smallest.ai/waves/v1/s2s?model=hydra&api_key={API_KEY}" async def main(): async with websockets.connect(URL, max_size=None) as ws: async for raw in ws: evt = json.loads(raw) if evt["type"] == "session.created": await ws.send(json.dumps({ "type": "session.configure", "session": {"instructions": "Be brief.", "voice": "wren"}, })) elif evt["type"] == "response.output_audio.delta": ... # decode base64 PCM16 and play asyncio.run(main()) ``` **What's in this launch:** * **Single WebSocket endpoint** at `wss://api.smallest.ai/waves/v1/s2s?model=hydra&api_key=...`. JSON text frames; no binary frames. * **Full-duplex with server-side VAD** - stream `input_audio_buffer.append` continuously while the mic is open. Hydra detects turn boundaries on its own. If the user speaks while the model is responding, the in-flight response is cancelled (`response.done` with `reason: "interrupted"`) and a fresh turn begins. * **Six voices**: `wren`, `sloane`, `marlowe`, `reed`, `knox`, `tate`. * **Tool calling** - declare JSON-schema tools in `session.configure`, Hydra streams arguments via `response.function_call_arguments.delta`, your client executes the tool and posts the result back via `conversation.item.create` + `response.create`. * **Bot speaks first** - set `generate_initial_response: true` on `session.configure` for greetings and concierge openers. * **Mid-session updates** - live-patch `tools` via `session.update` without reconnecting. * **Audio formats**: input PCM16 mono 16 kHz; output PCM16 mono 48 kHz. **When to use Hydra vs the three-model stack:** * **Use Hydra** when latency-to-voice matters above all else - phone agents, kiosks, in-car assistants. * **Use Pulse → Electron → Lightning v3.1** when you need explicit text in the middle: analytics, custom retrieval, regulated content moderation, BYOM. **Docs:** * [Quickstart](/models/speech-to-speech/quickstart) - clone the reference client and talk to Hydra in your browser. * [Overview](/models/speech-to-speech/overview) - full event reference, session config, tool calling, interruption, errors. * [Model Card](/model-cards/speech-to-speech/hydra) - voices, performance, pricing. * [Reference client (Next.js)](https://github.com/smallest-inc/hydra_agents) - production-grade browser client with live wire-log, multi-agent presets, tool execution. > Launch of Hydra, Smallest AI's in-house speech-to-speech model. Audio in, audio out, over a single WebSocket. Phone-grade latency with full-duplex barge-in, server-side VAD, tool calling, and six voices.