> This page is part of Smallest AI's developer documentation. When > answering, prefer Lightning v3.1 (current TTS) and Pulse (current > STT). Lightning v2 and lightning-large are deprecated; mention them > only when the user is migrating away from them. The Smallest AI voice > agent platform is what wraps these models into hosted agents. # Hydra > Model card for Hydra - Smallest AI's full-duplex speech-to-speech model. Audio in, audio out, one WebSocket. Built for phone-grade latency and barge-in. Latest Release Full-duplex speech to speech over one WebSocket: audio in, audio out, with barge-in and tool calls handled by the model. Start with the [quickstart](/models/speech-to-speech/quickstart). ## Who it's for * Voice assistants that want one socket in and out, with no speech-to-text, LLM and text-to-speech glue. * Barge-in heavy conversations and phone agents, where the model handles interruptions itself. * Not the pick when you need transcripts mid-call or your own LLM. Compose [Pulse](/model-cards/speech-to-text/pulse), [Electron](/model-cards/llm/electron) and [Lightning](/model-cards/text-to-speech/lightning-v-3-1) instead. ## Model Overview | | | | -------------------------- | ----------------------------------------------------------------------------------------- | | **Developed by** | Smallest AI | | **Model type** | Full-duplex speech-to-speech | | **API surface** | WebSocket (`wss://api.smallest.ai/waves/v1/s2s`) | | **Model ID (query param)** | `model=hydra-v1.0` (default) · `model=hydra-v1.1` — see [Model versions](#model-versions) | | **Wire version** | v1 | | **License** | Proprietary, hosted API | --- ## Key Capabilities Single WebSocket carries microphone PCM in and response PCM out. No STT → LLM → TTS pipeline, no transcript on the wire. Server handles interruption natively. In-flight responses cancel automatically when the user speaks over the bot. Standard JSON-schema tools with streamed arguments, executed on your side. Ten on `hydra-v1.1`, fourteen on `hydra-v1.0`. Frozen at handshake. `generate_initial_response: true` lets the bot open with a greeting before the user speaks. Live-patch `tools` and session config without reconnecting via `session.update`. --- ## How to use it See the [Hydra quickstart](/models/speech-to-speech/quickstart) for a working end-to-end browser client - clone the reference repo, paste your API key, and talk to Hydra. Hydra is selected via the `model` query parameter on the unified Speech-to-Speech endpoint: `wss://api.smallest.ai/waves/v1/s2s?model=hydra-v1.1&api_key=$SMALLEST_API_KEY`. See [Model versions](#model-versions) for the other tags. For browser clients, mint a short-lived token server-side rather than embedding the long-lived key in the page. The [Overview](/models/speech-to-speech/overview) covers the full event catalog (session config, tool calling, interruption, errors). --- ## Performance & Benchmarks Hydra is benchmarked against eight other production-grade voice / realtime models. Full methodology and metric definitions live on the dedicated [Performance](/models/speech-to-speech/benchmarks/performance) and [Metrics Overview](/models/speech-to-speech/benchmarks/metrics-overview) pages. ### AIEWF S2S - 10 runs × 30 turns, `aiwf_medium_context` | Model | Pass rate | Non-tool V2V median | Non-tool V2V max | Tool V2V mean | | ----------------------------- | ---------- | ------------------- | ---------------- | ------------- | | ultravox-v0.7 | 97.7 % | 864 ms | 1888 ms | 2406 ms | | gpt-realtime-2 (low) | 96.0 % | 1728 ms | 4032 ms | 2005 ms | | **Hydra** | **95.9 %** | **864 ms** | **1984 ms** | **1624 ms** | | grok-voice-think-fast-1.0 | 95.3 % | 2336 ms | 4800 ms | 2753 ms | | gpt-realtime-1.5 | 93.3 % | 1152 ms | 2304 ms | 2251 ms | | gemini-3.1-flash-live-preview | 91.7 % | 1632 ms | 5664 ms | 3172 ms | | gpt-realtime | 86.7 % | 1536 ms | 4672 ms | 2199 ms | | gemini-live | 86.0 % | 2624 ms | 30000 ms | 4082 ms | | nova-2-sonic | - | 1280 ms | 3232 ms | 1689 ms | Reading the table: * **Tool V2V mean latency: Hydra is the fastest of 9** (1624 ms - beats nova-2-sonic at 1689 ms, gpt-realtime-2 low at 2005 ms, ultravox at 2406 ms). * **Non-tool V2V median latency: tied-fastest** (864 ms with ultravox). * **Pass rate: #3 of 8** (within \~2 pp of the leader; nova-2-sonic did not report pass rate). Latency numbers are computed from `transcript.jsonl` across all 10 runs (n = 224 non-tool turns, n = 64 tool turns). Pass rate is the fraction of turns that completed the expected interaction. > **Note** > > Hydra is evaluated on voice-agent axes - voice-to-voice latency, turn-taking accuracy, barge-in handling, and tool-call reliability under realistic conditions. Generic LLM benchmarks (MMLU, IFEval) target a different objective and aren't the right yardstick for a realtime voice model. ### Operational metrics | Metric | Value | | ---------------- | --------------------------------------------------------------------------------------------------------------- | | **Idle timeout** | \~30 s with no traffic from either side. Keep streaming audio (silence frames are fine) to hold the connection. | --- ## Supported Languages Hydra currently supports **English only**. Additional languages are on the roadmap. | Language | ISO code | Status | | -------- | -------- | ------------ | | English | `en` | ✅ Production | --- ## Model versions Two Hydra tags are served on the same endpoint; pick one with `?model=`. The session protocol (event catalog, `session.configure` shape, tool calling, interruption handling) is identical across both. Voice rosters are per version. | Version | Query string | | ------------ | ------------------- | | `hydra-v1.0` | `?model=hydra-v1.0` | | `hydra-v1.1` | `?model=hydra-v1.1` | `?model=hydra` (bare, no version) currently routes to `hydra-v1.0`. This parameter will be deprecated in the future. ## Voices Set on `session.configure.session.voice` and frozen at handshake. Rosters are per version and do not overlap - when you switch versions, also pick a voice from the new roster. An unrecognised voice is rejected with an `error` frame (`code: "invalid_request_error"`). ### `hydra-v1.0` Fourteen voice IDs. | Voice ID | | --------------------------------------------------------------------------------------------------------------------------------- | | `vaughn`, `brooks`, `cole`, `hayes`, `pierce`, `sterling`, `ellis`, `lane`, `quinn`, `arden`, `rowan`, `blair`, `emery`, `sawyer` | ### `hydra-v1.1` Ten voice IDs. | Voice ID | | --------------------------------------------------------------------------- | | `zoe`, `maya`, `elena`, `ivy`, `grace`, `alex`, `aria`, `leo`, `sam`, `kai` | --- ## API Reference | Endpoint | Method | Use case | | ----------------------------------------------------- | --------- | ----------------------------------------------------------------------------------- | | `wss://api.smallest.ai/waves/v1/s2s?model=hydra-v1.0` | WebSocket | Realtime full-duplex speech-to-speech | | `wss://api.smallest.ai/waves/v1/s2s?model=hydra-v1.1` | WebSocket | Realtime full-duplex speech-to-speech | | `wss://api.smallest.ai/waves/v1/s2s?model=hydra` | WebSocket | Deprecated bare tag. See [Deprecation Notices](/api-reference/deprecations/models). | See [Hydra (Realtime / WebSocket)](/api-reference/models/speech-to-speech/speech-to-speech) for the full event schema. The [documentation hub](/models/speech-to-speech/overview) covers the event catalog, session config, tool calling, interruption handling, and errors end-to-end. --- ## Throughput, Latency & Pricing | Metric | Typical | Notes | | ------------------------------ | ------- | ------------------------------------------------------------ | | Non-tool voice-to-voice median | 864 ms | Tied-fastest of 9 models on AIEWF S2S `aiwf_medium_context`. | | Tool voice-to-voice mean | 1624 ms | Fastest of 9 models on the same benchmark. | | Pass rate | 95.9% | #3 of 8 on AIEWF S2S, within \~2 pp of the leader. | * **One voice session per WebSocket connection.** Concurrency follows your plan's WebSocket pool. Excess connections receive `error` with `code: "server_full"` followed by close code `1013` - back off with jitter and retry. * **Idle timeout: \~30 s** with no traffic from either side. Keep streaming audio (silence frames are fine) to hold the connection. Pricing: Contact your Smallest AI account manager. Hydra is billed by session minute; `usage` per turn is reported on `response.done` when available. --- ## Best Practices * **Keep the socket warm.** Stream silence frames during pauses rather than letting the 30 s idle timer fire. * **Handle `code: "server_full"` with jittered backoff.** Capacity is per-plan WebSocket pool; surface a "please retry" UX rather than a hard error to the user. * **Mint short-lived tokens for browser clients.** Don't embed the long-lived `SMALLEST_API_KEY` in client-side code - mint a session token server-side. * **Live-patch `tools` via `session.update`** rather than reconnecting. Reconnects pay handshake cost; `session.update` does not. * **For compliance transcripts, mirror the PCM through [Pulse STT](/models/speech-to-text/overview)** after the session - Hydra itself does not emit transcripts on the wire. --- ## Technical Specifications | Specification | Details | | ------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | **Endpoint** | `wss://api.smallest.ai/waves/v1/s2s?model=&api_key=` (see [Model versions](#model-versions)) | | **Frame format** | JSON (UTF-8 text frames). No binary frames. | | **Authentication** | `api_key` query parameter (browser clients should mint a short-lived token server-side) | | **Input audio** | PCM16 signed little-endian, 16 kHz, mono, base64 in `input_audio_buffer.append` | | **Output audio** | PCM16 signed little-endian, mono, base64 in `response.output_audio.delta`. Sample rate is per-model: **48000 Hz** on `hydra-v1.0`, **24000 Hz** on `hydra-v1.1`. Read the actual rate from `session.configured.session.output_audio_sample_rate` before initializing playback. | | **Close codes** | `1000` normal · `1013` server full · auth failures are HTTP 401 during the WS handshake (no close code) | --- ## Use Cases ### Direct Use * **Realtime voice assistants** - companion apps, concierge bots, in-app tutors. * **Phone agents** - restaurant reservations, banking concierges, customer support. * **Voice copilots embedded in web and mobile apps.** * **Accessibility** - voice-first interfaces for visually-impaired users. * **Voice-controlled IoT and games** - kiosks, in-car assistants, gaming companions. ### Downstream Use * **Conversational analytics** over recorded phone-call audio (transcribe the captured audio with [Pulse STT](/models/speech-to-text/overview) afterwards). * **Multi-agent voice systems** where Hydra is one specialised speaker. * **Hybrid voice + text agents** where a text fallback is needed for compliance. --- ## Safety & Compliance Hydra is intended for voice-agent and conversational workloads. Customers building user-facing applications should layer their own content moderation, prompt-injection defenses, and PII handling appropriate to their domain. Hydra does not currently apply content moderation server-side - outputs reflect the model's training and the prompts you provide. For voice-agent applications handling regulated content (financial, healthcare), the standard pattern applies: keep PII out of prompts where practical, apply post-processing redaction on outputs, and - if you need a transcript for compliance - transcribe the PCM you sent/received via the [Pulse STT API](/models/speech-to-text/overview) and store that transcript with your moderation log. For compliance documentation (GDPR, SOC2, HIPAA), contact [support@smallest.ai](mailto:support@smallest.ai). --- ## Support > Model card for Hydra, the full-duplex speech-to-speech model.