Hydra - Speech to Speech
Hydra is Smallest AI’s realtime, full-duplex speech-to-speech model.
Hydra is a realtime speech-to-speech model. The client streams microphone audio over a WebSocket, the model returns synthesised speech in the same socket, and turn-taking is handled server-side. There is no transcript on the wire - audio bytes are the payload.
If you’ve used the OpenAI Realtime API, Hydra fills the same role on the Smallest AI stack.
Three Hydra tags are served: ?model=hydra-v1.1 (current release, recommended), ?model=hydra-v1.0 (served, not deprecated), and ?model=hydra (deprecated on 2026-08-20; sessions still open but the server emits a warning frame with code: "model_deprecated" on connect). Omitting model behaves like ?model=hydra. Use hydra-v1.1 on all new integrations. Full per-version voice reference: Hydra model card. See also: Deprecation Notices.
Common use cases
What’s on the wire
Two things to know up front:
- Stream audio continuously - no manual
commitorend-of-turn. Hydra detects turn boundaries on its own. - Full-duplex - the user can speak over the model. The in-flight response cancels automatically with
status: "cancelled",reason: "interrupted".
Next
Related
- Model card - Hydra - voices, performance, pricing
- Reference client (Next.js) - production-grade browser client with barge-in, multi-agent presets, tool execution