Skip to navigation

Speech-to-Speech: Hydra V1.1 released, `?model=hydra` deprecated

Hydra V1.1 is the current release of the realtime speech-to-speech model. New integrations should connect with ?model=hydra-v1.1:

wss://api.smallest.ai/waves/v1/s2s?model=hydra-v1.1&api_key=<SMALLEST_API_KEY>

The session protocol (event catalog, session.configure shape, tool calling, interruption handling) is identical to the original hydra, so switching is a one-parameter change on the query string.

?model=hydra is deprecated. Existing sessions still open, but the server emits a warning frame with code: "model_deprecated" immediately before session.created. Migrate to ?model=hydra-v1.1. See Deprecation Notices.

Voice rosters are documented per version on the Hydra model card. An unknown voice on session.configure is rejected with an error frame (code: "invalid_request_error") rather than silently defaulted, so validate client-side.

hydra-v1.0 is also served on the same endpoint; it is not deprecated, and the server names it as the migration target when a client connects with the deprecated ?model=hydra.

Full voice + version reference on the Hydra model card. Overview + docs pages (overview, WebSocket connection, Audio I/O, Managing sessions, Tool calling) point at ?model=hydra-v1.1 in every URL example and use a V1.1 voice in the sample payloads. The voice field in Managing sessions points to the model card, which is the single source of truth for per-version voice tables, and documents the rejection of unknown voices.


Pulse STT realtime overview: VAD Events, Endpointing, Age & Gender, and Emotion Detection cards

The Realtime features overview now surfaces four capabilities that had dedicated feature pages but were missing from the overview:

  • VAD Events. Acoustic speech_started / speech_ended frames interleaved with the transcription stream.
  • Endpointing. Tune the end-of-utterance silence timeout with eou_timeout_ms.
  • Age & Gender Detection. Optional per-speaker inference alongside the transcript.
  • Emotion Detection. Optional per-utterance classification alongside the transcript.