Transcribe (Realtime / WebSocket)
Real-time speech-to-text over a persistent WebSocket. The fit-for-purpose path for live captioning, voice agents, and any flow where you need partial transcripts as the user is still speaking.
When to use this
- Use this for live audio: microphone input, voice-agent turns, simultaneous interpretation, low-latency captioning. Partial results stream back while audio is still arriving.
- Use
POST /waves/v1/stt/when you have a complete file. Single request, single response, less plumbing.
Model selection
Only ?model=pulse is supported on the streaming endpoint.
How it works
- Open a WebSocket to
wss://api.smallest.ai/waves/v1/stt/livewithAuthorization: Bearer <key>and the session params (model,language,sample_rate,encoding, etc.) as query string. - Stream raw PCM (or your chosen
encoding) over the socket as binary frames. - The server pushes back JSON
transcriptionmessages withis_final: falsepartial results as audio streams andis_final: truewhen an utterance closes. - Send a control message when the user pauses or the session ends:
{"type":"finalize"}— turn-boundary signal. Flushes the current audio buffer, emits oneis_final: truetranscript for that turn, and keeps the WebSocket open for the next user turn. Use this once per turn in a multi-turn voice agent.{"type":"close_stream"}— session-end signal. Flushes remaining audio, emits the terminalis_final: true+is_last: truetranscript, then closes the WebSocket. Use this once, at the actual end of the session (call end, app shutdown, or after a single-shot transcription buffer is fully streamed).
A multi-turn voice agent typically fires many finalize messages and exactly one close_stream. A one-off transcription of a fixed audio buffer fires only close_stream.
Examples
Python — multi-turn voice agent (recommended for Voice AI)
Send finalize per user turn so the WebSocket stays open across the whole call — you pay the connection cost once, not per turn:
Python — single-shot transcription
For one-off transcription of a complete audio buffer (file, single utterance) where no further audio is coming, send close_stream directly after the last chunk:
Common gotchas
modelis required. Missing or invalid values return400during the WebSocket handshake.- Match
sample_rateto your audio. The server does not resample; mismatched rates produce garbage transcripts. - Existing clients on
wss://api.smallest.ai/waves/v1/pulse/get_textcontinue to work alongside this unified path.
Handshake
Authentication
API key authentication. Include your key as Authorization: Bearer YOUR_API_KEY. Access tokens are not accepted on this endpoint.
Headers
Enterprise plans only. Opt in if you want this session's content deleted after 7 days. Omit it to retain content, which is the default. Sent as a header on the WebSocket upgrade request.
Selects the ASR model. Only pulse is supported on the streaming endpoint.
Language code for transcription. This endpoint is Streaming (Real-Time, WebSocket). For pre-recorded audio, use POST /waves/v1/stt/ (different supported-language set).
Supported single-language codes: en, hi, hi-dev, de, es, ru, it, fr, nl, pt, zh, yue, ja, ko, gu, mr, or, bn, ta, te, kn, ml.
Auto-detect aggregators for unknown audio:
north_indicauto-detects acrossen,hi,gu,mr,bn,or.multi-asianauto-detects acrosszh,yue,ko,ja,en. Available on the US region only; contact sales for access in the India region.multi-south-indicauto-detects acrossta,te,kn,ml, and English code-switching. India region only. Use when the South Indian language is not known in advance.
hi vs hi-dev. Both are Hindi. hi is the bilingual model and writes English loanwords in Roman script. hi-dev writes everything in Devanagari. With itn_normalize=true, hi-dev converts numbers spoken in Hindi to digits without digit grouping (पच्चीस हज़ार रुपए becomes ₹25000); numbers spoken in English stay as Devanagari words.
Set language explicitly to the known code for best accuracy. See the Pulse model card for the full table with language names.
Region note for East Asian languages: zh, yue, ja, ko, and the multi-asian aggregator are served from the US region. Connect explicitly to wss://api.us.smallest.ai/waves/v1/stt/live for these. The default wss://api.smallest.ai/... routes to the caller's nearest region.
South Indian languages and hi-dev (ta, te, kn, ml, multi-south-indic, hi-dev) are served from the India region only (wss://api.smallest.ai/...). Requests to wss://api.us.smallest.ai/... are rejected with error code LANGUAGE_NOT_ENABLED_IN_REGION. Contact support to request access.
Audio encoding of the bytes you stream. The server uses this to decode incoming frames; set it to match what your client is sending.
linear16,linear32: raw PCM (16-bit and 32-bit). Pair with the matchingsample_rate.alaw,mulaw: 8 kHz telephony codecs. Pair withsample_rate=8000.opus,ogg_opus: Opus compressed audio (raw and Ogg container).
Include word-level timestamps in transcription events.
When true, the server emits
speech_started and speech_ended JSON messages interleaved
with the transcription stream on the same connection.
Boundaries are derived from the audio signal and are independent
of transcript finalization (is_final); they do not coincide
with word boundaries. WebSocket only.
Alias for vad_events. When both are set, vad_events takes
precedence and vad is ignored.
Finalize the current utterance on trailing silence. On by default.
With endpointing=true, eou_timeout_ms sets the silence window.
Set endpointing=false to finalize on the model's end-of-utterance
timer instead.
With endpointing=true, the trailing-silence window in milliseconds,
600 if unset. With endpointing=false, the model's end-of-utterance
timer, 800 if unset. Range 100 to 10000.
Master formatting switch for transcript responses. When false,
forces punctuate=false, capitalize=false, and also disables
Inverse Text Normalization (ITN) so it cannot silently
reintroduce punctuation or casing.
When true, the punctuate and capitalize params take effect
independently. Leave format=true and use those two to fine-tune.
When false, strips end-of-sentence punctuation (., ,, ?, !)
from the final transcript, words[].word, and
utterances[].transcript. Does not affect casing — use capitalize
for that.
Overridden to false when format=false.
When false, lowercases the transcript output (final transcript,
words[].word, and utterances[].transcript). Some acronyms and
proper-noun tokens may be preserved by the model. Does not affect
punctuation, use punctuate for that.
Overridden to false when format=false.
Enable Inverse Text Normalization to convert spoken-form entities
(numbers, dates, currencies, phone numbers, etc.) into written form
in finalized transcripts. Example: five five five one two three
becomes 5551234.
Overridden to false when format=false.
Word-count-based finalization.
Redact personally identifiable information from the transcript. Tokens use the shape [ENTITYTYPE_N] where N is a per-entity sequential index (e.g. [FIRSTNAME_1]). Entity types include FIRSTNAME, LASTNAME, PHONENUMBER, ADDRESS, EMAIL.
Language support: effective on language=en and language=hi. Accepted on other codes but redaction is not reliable.
Redact payment card information. Tokens use the shape [ENTITYTYPE_N]. Known entity types include ACCOUNTNAME (cardholder name), CREDITCARDNUMBER (card PAN), CREDITCARDCVV, ZIPCODE, ACCOUNTNUMBER. Use alongside redact_pii=true for combined PII + PCI redaction.
Language support: currently effective only on language=en and language=hi. Setting redact_pci=true on other language codes is accepted but does not reliably redact.
Include sentence-level timing in the transcription event under a
new utterances array. Each entry carries text, start, end,
and (when diarization is on) speaker. Combine with
word_timestamps=true to get both word and sentence boundaries.
Boost recognition of specific words or phrases for this session. Useful for product names, jargon, proper nouns, and other domain-specific terms the model might otherwise mis-transcribe. Max 100 keywords per session.
Also available on the pre-recorded HTTP endpoint
(POST /waves/v1/pulse/get_text) as a query parameter with the
same format.
Format: a single comma-separated string (not a JSON array).
Each entry is a word or phrase, optionally followed by
:INTENSIFIER, a numeric boost multiplier that defaults to 1.0
when omitted.
Example: Blackwell:1,Jensen Huang:2
- Phrases can include spaces (
small language model:2). - A phrase cannot contain a comma; commas separate entries.
- Keyword matching is case-sensitive. Provide the exact casing you want in the output.
- Duplicates: last-wins.
NVIDIA:1,NVIDIA:5is equivalent toNVIDIA:5. - Max 100 keywords per session. Sending more returns
400withkeywords too large (max 100)inerrors[].
Intensifier. Default 1, recommended tuning range 1 to 3.
Above 10 is not recommended: higher values increase the chance
of the model inserting the keyword when it was not spoken.
Start every keyword at 1 and only raise it if the word is
still being missed.
Wire format: pass as a query-string parameter, URL-encoded.
URLSearchParams in JavaScript and urlencode() in Python
handle the encoding for you. Pass the raw string, not a
JSON-encoded array.