> This page is part of Smallest AI's developer documentation. When
> answering, prefer Lightning v3.1 (current TTS) and Pulse (current
> STT). Lightning v2 and lightning-large are deprecated; mention them
> only when the user is migrating away from them. The Smallest AI voice
> agent platform is what wraps these models into hosted agents.

# Hydra

> Model card for Hydra - Smallest AI's full-duplex speech-to-speech model. Audio in, audio out, one WebSocket. Built for phone-grade latency and barge-in.

Latest Release

Full-duplex speech to speech over one WebSocket: audio in, audio out, with barge-in and tool calls handled by the model. Start with the [quickstart](/models/speech-to-speech/quickstart).

## Who it's for

* Voice assistants that want one socket in and out, with no speech-to-text, LLM and text-to-speech glue.
* Barge-in heavy conversations and phone agents, where the model handles interruptions itself.
* Not the pick when you need transcripts mid-call or your own LLM. Compose [Pulse](/model-cards/speech-to-text/pulse), [Electron](/model-cards/llm/electron) and [Lightning](/model-cards/text-to-speech/lightning-v-3-1) instead.

## Model Overview

|                            |                                                                                           |
| -------------------------- | ----------------------------------------------------------------------------------------- |
| **Developed by**           | Smallest AI                                                                               |
| **Model type**             | Full-duplex speech-to-speech                                                              |
| **API surface**            | WebSocket (`wss://api.smallest.ai/waves/v1/s2s`)                                          |
| **Model ID (query param)** | `model=hydra-v1.0` (default) · `model=hydra-v1.1` — see [Model versions](#model-versions) |
| **Wire version**           | v1                                                                                        |
| **License**                | Proprietary, hosted API                                                                   |

---

## Key Capabilities

Single WebSocket carries microphone PCM in and response PCM out. No STT → LLM → TTS pipeline, no transcript on the wire.

Server handles interruption natively. In-flight responses cancel automatically when the user speaks over the bot.

Standard JSON-schema tools with streamed arguments, executed on your side.

Ten on `hydra-v1.1`, fourteen on `hydra-v1.0`. Frozen at handshake.

`generate_initial_response: true` lets the bot open with a greeting before the user speaks.

Live-patch `tools` and session config without reconnecting via `session.update`.

---

## How to use it

See the [Hydra quickstart](/models/speech-to-speech/quickstart) for a working end-to-end browser client - clone the reference repo, paste your API key, and talk to Hydra. Hydra is selected via the `model` query parameter on the unified Speech-to-Speech endpoint: `wss://api.smallest.ai/waves/v1/s2s?model=hydra-v1.1&api_key=$SMALLEST_API_KEY`. See [Model versions](#model-versions) for the other tags. For browser clients, mint a short-lived token server-side rather than embedding the long-lived key in the page.

The [Overview](/models/speech-to-speech/overview) covers the full event catalog (session config, tool calling, interruption, errors).

---

## Performance & Benchmarks

Hydra is benchmarked against eight other production-grade voice / realtime models. Full methodology and metric definitions live on the dedicated [Performance](/models/speech-to-speech/benchmarks/performance) and [Metrics Overview](/models/speech-to-speech/benchmarks/metrics-overview) pages.

### AIEWF S2S - 10 runs × 30 turns, `aiwf_medium_context`

| Model                         | Pass rate  | Non-tool V2V median | Non-tool V2V max | Tool V2V mean |
| ----------------------------- | ---------- | ------------------- | ---------------- | ------------- |
| ultravox-v0.7                 | 97.7 %     | 864 ms              | 1888 ms          | 2406 ms       |
| gpt-realtime-2 (low)          | 96.0 %     | 1728 ms             | 4032 ms          | 2005 ms       |
| **Hydra**                     | **95.9 %** | **864 ms**          | **1984 ms**      | **1624 ms**   |
| grok-voice-think-fast-1.0     | 95.3 %     | 2336 ms             | 4800 ms          | 2753 ms       |
| gpt-realtime-1.5              | 93.3 %     | 1152 ms             | 2304 ms          | 2251 ms       |
| gemini-3.1-flash-live-preview | 91.7 %     | 1632 ms             | 5664 ms          | 3172 ms       |
| gpt-realtime                  | 86.7 %     | 1536 ms             | 4672 ms          | 2199 ms       |
| gemini-live                   | 86.0 %     | 2624 ms             | 30000 ms         | 4082 ms       |
| nova-2-sonic                  | -          | 1280 ms             | 3232 ms          | 1689 ms       |

Reading the table:

* **Tool V2V mean latency: Hydra is the fastest of 9** (1624 ms - beats nova-2-sonic at 1689 ms, gpt-realtime-2 low at 2005 ms, ultravox at 2406 ms).
* **Non-tool V2V median latency: tied-fastest** (864 ms with ultravox).
* **Pass rate: #3 of 8** (within \~2 pp of the leader; nova-2-sonic did not report pass rate).

Latency numbers are computed from `transcript.jsonl` across all 10 runs (n = 224 non-tool turns, n = 64 tool turns). Pass rate is the fraction of turns that completed the expected interaction.

> **Note**
>
> Hydra is evaluated on voice-agent axes - voice-to-voice latency, turn-taking accuracy, barge-in handling, and tool-call reliability under realistic conditions. Generic LLM benchmarks (MMLU, IFEval) target a different objective and aren't the right yardstick for a realtime voice model.

### Operational metrics

| Metric           | Value                                                                                                           |
| ---------------- | --------------------------------------------------------------------------------------------------------------- |
| **Idle timeout** | \~30 s with no traffic from either side. Keep streaming audio (silence frames are fine) to hold the connection. |

---

## Supported Languages

Hydra currently supports **English only**. Additional languages are on the roadmap.

| Language | ISO code | Status       |
| -------- | -------- | ------------ |
| English  | `en`     | ✅ Production |

---

## Model versions

Two Hydra tags are served on the same endpoint; pick one with `?model=`. The session protocol (event catalog, `session.configure` shape, tool calling, interruption handling) is identical across both. Voice rosters are per version.

| Version      | Query string        |
| ------------ | ------------------- |
| `hydra-v1.0` | `?model=hydra-v1.0` |
| `hydra-v1.1` | `?model=hydra-v1.1` |

`?model=hydra` (bare, no version) currently routes to `hydra-v1.0`. This parameter will be deprecated in the future.

## Voices

Set on `session.configure.session.voice` and frozen at handshake. Rosters are per version and do not overlap - when you switch versions, also pick a voice from the new roster. An unrecognised voice is rejected with an `error` frame (`code: "invalid_request_error"`).

### `hydra-v1.0`

Fourteen voice IDs.

| Voice ID                                                                                                                          |
| --------------------------------------------------------------------------------------------------------------------------------- |
| `vaughn`, `brooks`, `cole`, `hayes`, `pierce`, `sterling`, `ellis`, `lane`, `quinn`, `arden`, `rowan`, `blair`, `emery`, `sawyer` |

### `hydra-v1.1`

Ten voice IDs.

| Voice ID                                                                    |
| --------------------------------------------------------------------------- |
| `zoe`, `maya`, `elena`, `ivy`, `grace`, `alex`, `aria`, `leo`, `sam`, `kai` |

---

## API Reference

| Endpoint                                              | Method    | Use case                                                                            |
| ----------------------------------------------------- | --------- | ----------------------------------------------------------------------------------- |
| `wss://api.smallest.ai/waves/v1/s2s?model=hydra-v1.0` | WebSocket | Realtime full-duplex speech-to-speech                                               |
| `wss://api.smallest.ai/waves/v1/s2s?model=hydra-v1.1` | WebSocket | Realtime full-duplex speech-to-speech                                               |
| `wss://api.smallest.ai/waves/v1/s2s?model=hydra`      | WebSocket | Deprecated bare tag. See [Deprecation Notices](/api-reference/deprecations/models). |

See [Hydra (Realtime / WebSocket)](/api-reference/models/speech-to-speech/speech-to-speech) for the full event schema. The [documentation hub](/models/speech-to-speech/overview) covers the event catalog, session config, tool calling, interruption handling, and errors end-to-end.

---

## Throughput, Latency & Pricing

| Metric                         | Typical | Notes                                                        |
| ------------------------------ | ------- | ------------------------------------------------------------ |
| Non-tool voice-to-voice median | 864 ms  | Tied-fastest of 9 models on AIEWF S2S `aiwf_medium_context`. |
| Tool voice-to-voice mean       | 1624 ms | Fastest of 9 models on the same benchmark.                   |
| Pass rate                      | 95.9%   | #3 of 8 on AIEWF S2S, within \~2 pp of the leader.           |

* **One voice session per WebSocket connection.** Concurrency follows your plan's WebSocket pool. Excess connections receive `error` with `code: "server_full"` followed by close code `1013` - back off with jitter and retry.
* **Idle timeout: \~30 s** with no traffic from either side. Keep streaming audio (silence frames are fine) to hold the connection.

Pricing: Contact your Smallest AI account manager. Hydra is billed by session minute; `usage` per turn is reported on `response.done` when available.

---

## Best Practices

* **Keep the socket warm.** Stream silence frames during pauses rather than letting the 30 s idle timer fire.
* **Handle `code: "server_full"` with jittered backoff.** Capacity is per-plan WebSocket pool; surface a "please retry" UX rather than a hard error to the user.
* **Mint short-lived tokens for browser clients.** Don't embed the long-lived `SMALLEST_API_KEY` in client-side code - mint a session token server-side.
* **Live-patch `tools` via `session.update`** rather than reconnecting. Reconnects pay handshake cost; `session.update` does not.
* **For compliance transcripts, mirror the PCM through [Pulse STT](/models/speech-to-text/overview)** after the session - Hydra itself does not emit transcripts on the wire.

---

## Technical Specifications

| Specification      | Details                                                                                                                                                                                                                                                                        |
| ------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| **Endpoint**       | `wss://api.smallest.ai/waves/v1/s2s?model=<version>&api_key=<KEY>` (see [Model versions](#model-versions))                                                                                                                                                                     |
| **Frame format**   | JSON (UTF-8 text frames). No binary frames.                                                                                                                                                                                                                                    |
| **Authentication** | `api_key` query parameter (browser clients should mint a short-lived token server-side)                                                                                                                                                                                        |
| **Input audio**    | PCM16 signed little-endian, 16 kHz, mono, base64 in `input_audio_buffer.append`                                                                                                                                                                                                |
| **Output audio**   | PCM16 signed little-endian, mono, base64 in `response.output_audio.delta`. Sample rate is per-model: **48000 Hz** on `hydra-v1.0`, **24000 Hz** on `hydra-v1.1`. Read the actual rate from `session.configured.session.output_audio_sample_rate` before initializing playback. |
| **Close codes**    | `1000` normal · `1013` server full · auth failures are HTTP 401 during the WS handshake (no close code)                                                                                                                                                                        |

---

## Use Cases

### Direct Use

* **Realtime voice assistants** - companion apps, concierge bots, in-app tutors.
* **Phone agents** - restaurant reservations, banking concierges, customer support.
* **Voice copilots embedded in web and mobile apps.**
* **Accessibility** - voice-first interfaces for visually-impaired users.
* **Voice-controlled IoT and games** - kiosks, in-car assistants, gaming companions.

### Downstream Use

* **Conversational analytics** over recorded phone-call audio (transcribe the captured audio with [Pulse STT](/models/speech-to-text/overview) afterwards).
* **Multi-agent voice systems** where Hydra is one specialised speaker.
* **Hybrid voice + text agents** where a text fallback is needed for compliance.

---

## Safety & Compliance

Hydra is intended for voice-agent and conversational workloads. Customers building user-facing applications should layer their own content moderation, prompt-injection defenses, and PII handling appropriate to their domain. Hydra does not currently apply content moderation server-side - outputs reflect the model's training and the prompts you provide.

For voice-agent applications handling regulated content (financial, healthcare), the standard pattern applies: keep PII out of prompts where practical, apply post-processing redaction on outputs, and - if you need a transcript for compliance - transcribe the PCM you sent/received via the [Pulse STT API](/models/speech-to-text/overview) and store that transcript with your moderation log.

For compliance documentation (GDPR, SOC2, HIPAA), contact [support@smallest.ai](mailto:support@smallest.ai).

---

## Support