Hydra

View as Markdown
Latest Release

Hydra is Smallest AI’s speech-to-speech model. A single WebSocket carries microphone audio from your client to the model and streams synthesised response audio back. There is no STT → LLM → TTS pipeline in the middle, and no transcript on the wire.

Jump to: Benchmarks · Voices · API Reference · Pricing & Throughput · Quickstart

Model Overview

Developed bySmallest AI
Model typeFull-duplex speech-to-speech
API surfaceWebSocket (wss://api.smallest.ai/waves/v1/s2s)
Model ID (query param)model=hydra-v1.1 (default) · model=hydra-v1.0 · model=hydra (deprecated) - see Model versions
Wire versionv1
LicenseProprietary, hosted API

Key Capabilities

Audio In, Audio Out

Single WebSocket carries microphone PCM in and response PCM out. No STT → LLM → TTS pipeline, no transcript on the wire.

Full-duplex Barge-in

Server handles interruption natively. In-flight responses cancel automatically when the user speaks over the bot.

Tool / Function Calling

Standard JSON-schema tools with streamed arguments, executed on your side.

Per-version Voices

Ten on hydra-v1.1, fourteen on hydra-v1.0. Frozen at handshake.

Bot Speaks First

generate_initial_response: true lets the bot open with a greeting before the user speaks.

Mid-session Updates

Live-patch tools and session config without reconnecting via session.update.


How to use it

See the Hydra quickstart for a working end-to-end browser client - clone the reference repo, paste your API key, and talk to Hydra. Hydra is selected via the model query parameter on the unified Speech-to-Speech endpoint: wss://api.smallest.ai/waves/v1/s2s?model=hydra-v1.1&api_key=$SMALLEST_API_KEY. See Model versions for the other tags. For browser clients, mint a short-lived token server-side rather than embedding the long-lived key in the page.

The Overview covers the full event catalog (session config, tool calling, interruption, errors).


Performance & Benchmarks

Hydra is benchmarked against eight other production-grade voice / realtime models. Full methodology and metric definitions live on the dedicated Performance and Metrics Overview pages.

AIEWF S2S - 10 runs × 30 turns, aiwf_medium_context

ModelPass rateNon-tool V2V medianNon-tool V2V maxTool V2V mean
ultravox-v0.797.7 %864 ms1888 ms2406 ms
gpt-realtime-2 (low)96.0 %1728 ms4032 ms2005 ms
Hydra95.9 %864 ms1984 ms1624 ms
grok-voice-think-fast-1.095.3 %2336 ms4800 ms2753 ms
gpt-realtime-1.593.3 %1152 ms2304 ms2251 ms
gemini-3.1-flash-live-preview91.7 %1632 ms5664 ms3172 ms
gpt-realtime86.7 %1536 ms4672 ms2199 ms
gemini-live86.0 %2624 ms30000 ms4082 ms
nova-2-sonic-1280 ms3232 ms1689 ms

Reading the table:

  • Tool V2V mean latency: Hydra is the fastest of 9 (1624 ms - beats nova-2-sonic at 1689 ms, gpt-realtime-2 low at 2005 ms, ultravox at 2406 ms).
  • Non-tool V2V median latency: tied-fastest (864 ms with ultravox).
  • Pass rate: #3 of 8 (within ~2 pp of the leader; nova-2-sonic did not report pass rate).

Latency numbers are computed from transcript.jsonl across all 10 runs (n = 224 non-tool turns, n = 64 tool turns). Pass rate is the fraction of turns that completed the expected interaction.

Hydra is evaluated on voice-agent axes - voice-to-voice latency, turn-taking accuracy, barge-in handling, and tool-call reliability under realistic conditions. Generic LLM benchmarks (MMLU, IFEval) target a different objective and aren’t the right yardstick for a realtime voice model.

Operational metrics

MetricValue
Idle timeout~30 s with no traffic from either side. Keep streaming audio (silence frames are fine) to hold the connection.

Supported Languages

Hydra currently supports English only. Additional languages are on the roadmap.

LanguageISO codeStatus
Englishen✅ Production

Model versions

Three Hydra tags are served on the same endpoint; pick one with ?model=. The session protocol (event catalog, session.configure shape, tool calling, interruption handling) is identical across all three, so switching versions is a one-parameter change on the query string.

VersionQuery stringNotes
hydra-v1.1 (current release, default)?model=hydra-v1.1Latest speech-to-speech model. Recommended for all new integrations.
hydra-v1.0?model=hydra-v1.0Served, not deprecated. The server’s deprecation message for ?model=hydra names this tag as the migration target.
hydra (original release) Deprecated?model=hydraOriginal release. Existing sessions still open; the server emits a warning frame with code: "model_deprecated" immediately before session.created. Listed on the Deprecation Notices page.

Voices

Set on session.configure.session.voice and frozen at handshake. Rosters are per version and do not overlap - when you switch versions, also pick a voice from the new roster. Omit voice and the server applies sterling. An unrecognised voice is rejected with an error frame (code: "invalid_request_error"), not silently defaulted, so validate client-side.

hydra-v1.1 (current release)

Ten voice IDs. This is the roster to use for new integrations.

Voice ID
zoe, maya, elena, ivy, grace, alex, aria, leo, sam, kai

hydra-v1.0

Fourteen voice IDs.

Voice ID
vaughn, brooks, cole, hayes, pierce, sterling, ellis, lane, quinn, arden, rowan, blair, emery, sawyer

hydra (original release) Deprecated

Migrate to hydra-v1.1 or hydra-v1.0 and pick a voice from that version’s roster above.


API Reference

EndpointMethodUse case
wss://api.smallest.ai/waves/v1/s2s?model=hydra-v1.1WebSocketRealtime full-duplex speech-to-speech (current release)
wss://api.smallest.ai/waves/v1/s2s?model=hydra-v1.0WebSocketRealtime full-duplex speech-to-speech
wss://api.smallest.ai/waves/v1/s2s?model=hydraWebSocketRealtime full-duplex speech-to-speech (original release, deprecated)

See Hydra (Realtime / WebSocket) for the full event schema. The documentation hub covers the event catalog, session config, tool calling, interruption handling, and errors end-to-end.


Throughput, Latency & Pricing

MetricTypicalNotes
Non-tool voice-to-voice median864 msTied-fastest of 9 models on AIEWF S2S aiwf_medium_context.
Tool voice-to-voice mean1624 msFastest of 9 models on the same benchmark.
Pass rate95.9%#3 of 8 on AIEWF S2S, within ~2 pp of the leader.
  • One voice session per WebSocket connection. Concurrency follows your plan’s WebSocket pool. Excess connections receive error with code: "server_full" followed by close code 1013 - back off with jitter and retry.
  • Idle timeout: ~30 s with no traffic from either side. Keep streaming audio (silence frames are fine) to hold the connection.

Pricing: Contact your Smallest AI account manager. Hydra is billed by session minute; usage per turn is reported on response.done when available.


Best Practices

  • Keep the socket warm. Stream silence frames during pauses rather than letting the 30 s idle timer fire.
  • Handle code: "server_full" with jittered backoff. Capacity is per-plan WebSocket pool; surface a “please retry” UX rather than a hard error to the user.
  • Mint short-lived tokens for browser clients. Don’t embed the long-lived SMALLEST_API_KEY in client-side code - mint a session token server-side.
  • Live-patch tools via session.update rather than reconnecting. Reconnects pay handshake cost; session.update does not.
  • For compliance transcripts, mirror the PCM through Pulse STT after the session - Hydra itself does not emit transcripts on the wire.

Technical Specifications

SpecificationDetails
Endpointwss://api.smallest.ai/waves/v1/s2s?model=hydra-v1.1&api_key=<KEY> (current release; see Model versions)
Frame formatJSON (UTF-8 text frames). No binary frames.
Authenticationapi_key query parameter (browser clients should mint a short-lived token server-side)
Input audioPCM16 signed little-endian, 16 kHz, mono, base64 in input_audio_buffer.append
Output audioPCM16 signed little-endian, 48 kHz, mono, base64 in response.output_audio.delta
Close codes1000 normal · 1013 server full · auth failures are HTTP 401 during the WS handshake (no close code)

Use Cases

Direct Use

  • Realtime voice assistants - companion apps, concierge bots, in-app tutors.
  • Phone agents - restaurant reservations, banking concierges, customer support.
  • Voice copilots embedded in web and mobile apps.
  • Accessibility - voice-first interfaces for visually-impaired users.
  • Voice-controlled IoT and games - kiosks, in-car assistants, gaming companions.

Downstream Use

  • Conversational analytics over recorded phone-call audio (transcribe the captured audio with Pulse STT afterwards).
  • Multi-agent voice systems where Hydra is one specialised speaker.
  • Hybrid voice + text agents where a text fallback is needed for compliance.

Safety & Compliance

Hydra is intended for voice-agent and conversational workloads. Customers building user-facing applications should layer their own content moderation, prompt-injection defenses, and PII handling appropriate to their domain. Hydra does not currently apply content moderation server-side - outputs reflect the model’s training and the prompts you provide.

For voice-agent applications handling regulated content (financial, healthcare), the standard pattern applies: keep PII out of prompts where practical, apply post-processing redaction on outputs, and - if you need a transcript for compliance - transcribe the PCM you sent/received via the Pulse STT API and store that transcript with your moderation log.

For compliance documentation (GDPR, SOC2, HIPAA), contact support@smallest.ai.


Support