Skip to navigation

Hydra

Model card for Hydra, the full-duplex speech-to-speech model.
View as Markdown
Latest Release

Full-duplex speech to speech over one WebSocket: audio in, audio out, with barge-in and tool calls handled by the model. Start with the quickstart.

Who it’s for

  • Voice assistants that want one socket in and out, with no speech-to-text, LLM and text-to-speech glue.
  • Barge-in heavy conversations and phone agents, where the model handles interruptions itself.
  • Not the pick when you need transcripts mid-call or your own LLM. Compose Pulse, Electron and Lightning instead.

Model Overview

Developed bySmallest AI
Model typeFull-duplex speech-to-speech
API surfaceWebSocket (wss://api.smallest.ai/waves/v1/s2s)
Model ID (query param)model=hydra-v1.0 (default) · model=hydra-v1.1 — see Model versions
Wire versionv1
LicenseProprietary, hosted API

Key Capabilities

Audio In, Audio Out

Single WebSocket carries microphone PCM in and response PCM out. No STT → LLM → TTS pipeline, no transcript on the wire.

Full-duplex Barge-in

Server handles interruption natively. In-flight responses cancel automatically when the user speaks over the bot.

Tool / Function Calling

Standard JSON-schema tools with streamed arguments, executed on your side.

Per-version Voices

Ten on hydra-v1.1, fourteen on hydra-v1.0. Frozen at handshake.

Bot Speaks First

generate_initial_response: true lets the bot open with a greeting before the user speaks.

Mid-session Updates

Live-patch tools and session config without reconnecting via session.update.


How to use it

See the Hydra quickstart for a working end-to-end browser client - clone the reference repo, paste your API key, and talk to Hydra. Hydra is selected via the model query parameter on the unified Speech-to-Speech endpoint: wss://api.smallest.ai/waves/v1/s2s?model=hydra-v1.1&api_key=$SMALLEST_API_KEY. See Model versions for the other tags. For browser clients, mint a short-lived token server-side rather than embedding the long-lived key in the page.

The Overview covers the full event catalog (session config, tool calling, interruption, errors).


Performance & Benchmarks

Hydra is benchmarked against eight other production-grade voice / realtime models. Full methodology and metric definitions live on the dedicated Performance and Metrics Overview pages.

AIEWF S2S - 10 runs × 30 turns, aiwf_medium_context

ModelPass rateNon-tool V2V medianNon-tool V2V maxTool V2V mean
ultravox-v0.797.7 %864 ms1888 ms2406 ms
gpt-realtime-2 (low)96.0 %1728 ms4032 ms2005 ms
Hydra95.9 %864 ms1984 ms1624 ms
grok-voice-think-fast-1.095.3 %2336 ms4800 ms2753 ms
gpt-realtime-1.593.3 %1152 ms2304 ms2251 ms
gemini-3.1-flash-live-preview91.7 %1632 ms5664 ms3172 ms
gpt-realtime86.7 %1536 ms4672 ms2199 ms
gemini-live86.0 %2624 ms30000 ms4082 ms
nova-2-sonic-1280 ms3232 ms1689 ms

Reading the table:

  • Tool V2V mean latency: Hydra is the fastest of 9 (1624 ms - beats nova-2-sonic at 1689 ms, gpt-realtime-2 low at 2005 ms, ultravox at 2406 ms).
  • Non-tool V2V median latency: tied-fastest (864 ms with ultravox).
  • Pass rate: #3 of 8 (within ~2 pp of the leader; nova-2-sonic did not report pass rate).

Latency numbers are computed from transcript.jsonl across all 10 runs (n = 224 non-tool turns, n = 64 tool turns). Pass rate is the fraction of turns that completed the expected interaction.

Hydra is evaluated on voice-agent axes - voice-to-voice latency, turn-taking accuracy, barge-in handling, and tool-call reliability under realistic conditions. Generic LLM benchmarks (MMLU, IFEval) target a different objective and aren’t the right yardstick for a realtime voice model.

Operational metrics

MetricValue
Idle timeout~30 s with no traffic from either side. Keep streaming audio (silence frames are fine) to hold the connection.

Supported Languages

Hydra currently supports English only. Additional languages are on the roadmap.

LanguageISO codeStatus
Englishen✅ Production

Model versions

Two Hydra tags are served on the same endpoint; pick one with ?model=. The session protocol (event catalog, session.configure shape, tool calling, interruption handling) is identical across both. Voice rosters are per version.

VersionQuery string
hydra-v1.0?model=hydra-v1.0
hydra-v1.1?model=hydra-v1.1

?model=hydra (bare, no version) currently routes to hydra-v1.0. This parameter will be deprecated in the future.

Voices

Set on session.configure.session.voice and frozen at handshake. Rosters are per version and do not overlap - when you switch versions, also pick a voice from the new roster. An unrecognised voice is rejected with an error frame (code: "invalid_request_error").

hydra-v1.0

Fourteen voice IDs.

Voice ID
vaughn, brooks, cole, hayes, pierce, sterling, ellis, lane, quinn, arden, rowan, blair, emery, sawyer

hydra-v1.1

Ten voice IDs.

Voice ID
zoe, maya, elena, ivy, grace, alex, aria, leo, sam, kai

API Reference

EndpointMethodUse case
wss://api.smallest.ai/waves/v1/s2s?model=hydra-v1.0WebSocketRealtime full-duplex speech-to-speech
wss://api.smallest.ai/waves/v1/s2s?model=hydra-v1.1WebSocketRealtime full-duplex speech-to-speech
wss://api.smallest.ai/waves/v1/s2s?model=hydraWebSocketDeprecated bare tag. See Deprecation Notices.

See Hydra (Realtime / WebSocket) for the full event schema. The documentation hub covers the event catalog, session config, tool calling, interruption handling, and errors end-to-end.


Throughput, Latency & Pricing

MetricTypicalNotes
Non-tool voice-to-voice median864 msTied-fastest of 9 models on AIEWF S2S aiwf_medium_context.
Tool voice-to-voice mean1624 msFastest of 9 models on the same benchmark.
Pass rate95.9%#3 of 8 on AIEWF S2S, within ~2 pp of the leader.
  • One voice session per WebSocket connection. Concurrency follows your plan’s WebSocket pool. Excess connections receive error with code: "server_full" followed by close code 1013 - back off with jitter and retry.
  • Idle timeout: ~30 s with no traffic from either side. Keep streaming audio (silence frames are fine) to hold the connection.

Pricing: Contact your Smallest AI account manager. Hydra is billed by session minute; usage per turn is reported on response.done when available.


Best Practices

  • Keep the socket warm. Stream silence frames during pauses rather than letting the 30 s idle timer fire.
  • Handle code: "server_full" with jittered backoff. Capacity is per-plan WebSocket pool; surface a “please retry” UX rather than a hard error to the user.
  • Mint short-lived tokens for browser clients. Don’t embed the long-lived SMALLEST_API_KEY in client-side code - mint a session token server-side.
  • Live-patch tools via session.update rather than reconnecting. Reconnects pay handshake cost; session.update does not.
  • For compliance transcripts, mirror the PCM through Pulse STT after the session - Hydra itself does not emit transcripts on the wire.

Technical Specifications

SpecificationDetails
Endpointwss://api.smallest.ai/waves/v1/s2s?model=<version>&api_key=<KEY> (see Model versions)
Frame formatJSON (UTF-8 text frames). No binary frames.
Authenticationapi_key query parameter (browser clients should mint a short-lived token server-side)
Input audioPCM16 signed little-endian, 16 kHz, mono, base64 in input_audio_buffer.append
Output audioPCM16 signed little-endian, mono, base64 in response.output_audio.delta. Sample rate is per-model: 48000 Hz on hydra-v1.0, 24000 Hz on hydra-v1.1. Read the actual rate from session.configured.session.output_audio_sample_rate before initializing playback.
Close codes1000 normal · 1013 server full · auth failures are HTTP 401 during the WS handshake (no close code)

Use Cases

Direct Use

  • Realtime voice assistants - companion apps, concierge bots, in-app tutors.
  • Phone agents - restaurant reservations, banking concierges, customer support.
  • Voice copilots embedded in web and mobile apps.
  • Accessibility - voice-first interfaces for visually-impaired users.
  • Voice-controlled IoT and games - kiosks, in-car assistants, gaming companions.

Downstream Use

  • Conversational analytics over recorded phone-call audio (transcribe the captured audio with Pulse STT afterwards).
  • Multi-agent voice systems where Hydra is one specialised speaker.
  • Hybrid voice + text agents where a text fallback is needed for compliance.

Safety & Compliance

Hydra is intended for voice-agent and conversational workloads. Customers building user-facing applications should layer their own content moderation, prompt-injection defenses, and PII handling appropriate to their domain. Hydra does not currently apply content moderation server-side - outputs reflect the model’s training and the prompts you provide.

For voice-agent applications handling regulated content (financial, healthcare), the standard pattern applies: keep PII out of prompts where practical, apply post-processing redaction on outputs, and - if you need a transcript for compliance - transcribe the PCM you sent/received via the Pulse STT API and store that transcript with your moderation log.

For compliance documentation (GDPR, SOC2, HIPAA), contact support@smallest.ai.


Support