Skip to navigation

Transcribe (Realtime / WebSocket)

View as Markdown

Real-time speech-to-text over a persistent WebSocket. The fit-for-purpose path for live captioning, voice agents, and any flow where you need partial transcripts as the user is still speaking.

When to use this

  • Use this for live audio: microphone input, voice-agent turns, simultaneous interpretation, low-latency captioning. Partial results stream back while audio is still arriving.
  • Use POST /waves/v1/stt/ when you have a complete file. Single request, single response, less plumbing.

Model selection

Only ?model=pulse is supported on the streaming endpoint.

How it works

  1. Open a WebSocket to wss://api.smallest.ai/waves/v1/stt/live with Authorization: Bearer <key> and the session params (model, language, sample_rate, encoding, etc.) as query string.
  2. Stream raw PCM (or your chosen encoding) over the socket as binary frames.
  3. The server pushes back JSON transcription messages with is_final: false partial results as audio streams and is_final: true when an utterance closes.
  4. Send a control message when the user pauses or the session ends:
    • {"type":"finalize"} — turn-boundary signal. Flushes the current audio buffer, emits one is_final: true transcript for that turn, and keeps the WebSocket open for the next user turn. Use this once per turn in a multi-turn voice agent.
    • {"type":"close_stream"} — session-end signal. Flushes remaining audio, emits the terminal is_final: true + is_last: true transcript, then closes the WebSocket. Use this once, at the actual end of the session (call end, app shutdown, or after a single-shot transcription buffer is fully streamed).

A multi-turn voice agent typically fires many finalize messages and exactly one close_stream. A one-off transcription of a fixed audio buffer fires only close_stream.

Examples

Python — multi-turn voice agent (recommended for Voice AI)

Send finalize per user turn so the WebSocket stays open across the whole call — you pay the connection cost once, not per turn:

import asyncio, json, websockets
URL = "wss://api.smallest.ai/waves/v1/stt/live?model=pulse&language=en&sample_rate=16000&encoding=linear16&itn_normalize=true&eou_timeout_ms=1000"
HEADERS = {"Authorization": f"Bearer {API_KEY}"}
async def run_voice_agent(audio_source, llm_reply, stop_event):
async with websockets.connect(URL, additional_headers=HEADERS) as ws:
async def stream_audio():
async for frame in audio_source:
if stop_event.is_set(): return
await ws.send(frame)
# Call this when your VAD detects end-of-turn (user paused)
async def end_of_turn():
await ws.send(json.dumps({"type": "finalize"}))
async def consume():
async for msg in ws:
data = json.loads(msg)
if data.get("is_last"): break # only fires after close_stream
if data.get("is_final"):
await llm_reply(data["transcript"]) # ITN-normalized full turn
producer = asyncio.create_task(stream_audio())
consumer = asyncio.create_task(consume())
await stop_event.wait() # end of call
await ws.send(json.dumps({"type": "close_stream"}))
await consumer
producer.cancel()

Python — single-shot transcription

For one-off transcription of a complete audio buffer (file, single utterance) where no further audio is coming, send close_stream directly after the last chunk:

import asyncio, json, websockets
URL = "wss://api.smallest.ai/waves/v1/stt/live?model=pulse&language=en&sample_rate=16000&encoding=linear16"
HEADERS = {"Authorization": f"Bearer {API_KEY}"}
async def transcribe_once(audio_bytes):
async with websockets.connect(URL, additional_headers=HEADERS) as ws:
for i in range(0, len(audio_bytes), 4096):
await ws.send(audio_bytes[i:i+4096])
await ws.send(json.dumps({"type": "close_stream"}))
async for msg in ws:
data = json.loads(msg)
if data.get("is_final"):
print(data["transcript"])
if data.get("is_last"):
break
asyncio.run(transcribe_once(open("audio.pcm", "rb").read()))

Common gotchas

  • model is required. Missing or invalid values return 400 during the WebSocket handshake.
  • Match sample_rate to your audio. The server does not resample; mismatched rates produce garbage transcripts.
  • Existing clients on wss://api.smallest.ai/waves/v1/pulse/get_text continue to work alongside this unified path.

Handshake

WSS
wss://api.smallest.ai/waves/v1/stt/live

Authentication

AuthorizationBearer

API key authentication. Include your key as Authorization: Bearer YOUR_API_KEY. Access tokens are not accepted on this endpoint.

Headers

x-expire-contentstringRequired

Enterprise plans only. Opt in if you want this session's content deleted after 7 days. Omit it to retain content, which is the default. Sent as a header on the WebSocket upgrade request.

modelanyRequired

Selects the ASR model. Only pulse is supported on the streaming endpoint.

languageanyOptional

Language code for transcription. This endpoint is Streaming (Real-Time, WebSocket). For pre-recorded audio, use POST /waves/v1/stt/ (different supported-language set).

Supported single-language codes: en, hi, hi-dev, de, es, ru, it, fr, nl, pt, zh, yue, ja, ko, gu, mr, or, bn, ta, te, kn, ml.

Auto-detect aggregators for unknown audio:

  • north_indic auto-detects across en, hi, gu, mr, bn, or.
  • multi-asian auto-detects across zh, yue, ko, ja, en. Available on the US region only; contact sales for access in the India region.
  • multi-south-indic auto-detects across ta, te, kn, ml, and English code-switching. India region only. Use when the South Indian language is not known in advance.

hi vs hi-dev. Both are Hindi. hi is the bilingual model and writes English loanwords in Roman script. hi-dev writes everything in Devanagari. With itn_normalize=true, hi-dev converts numbers spoken in Hindi to digits without digit grouping (पच्चीस हज़ार रुपए becomes ₹25000); numbers spoken in English stay as Devanagari words.

Set language explicitly to the known code for best accuracy. See the Pulse model card for the full table with language names.

Region note for East Asian languages: zh, yue, ja, ko, and the multi-asian aggregator are served from the US region. Connect explicitly to wss://api.us.smallest.ai/waves/v1/stt/live for these. The default wss://api.smallest.ai/... routes to the caller's nearest region.

South Indian languages and hi-dev (ta, te, kn, ml, multi-south-indic, hi-dev) are served from the India region only (wss://api.smallest.ai/...). Requests to wss://api.us.smallest.ai/... are rejected with error code LANGUAGE_NOT_ENABLED_IN_REGION. Contact support to request access.

sample_rateanyOptionalDefaults to 16000
Audio sample rate in Hz of the bytes you stream. Must match the actual rate of your audio source.
encodinganyOptionalDefaults to linear16

Audio encoding of the bytes you stream. The server uses this to decode incoming frames; set it to match what your client is sending.

  • linear16, linear32: raw PCM (16-bit and 32-bit). Pair with the matching sample_rate.
  • alaw, mulaw: 8 kHz telephony codecs. Pair with sample_rate=8000.
  • opus, ogg_opus: Opus compressed audio (raw and Ogg container).
word_timestampsanyOptionalDefaults to false

Include word-level timestamps in transcription events.

diarizeanyOptionalDefaults to false
Enable speaker diarization to identify different speakers in the audio.
vad_eventsanyOptionalDefaults to false

When true, the server emits speech_started and speech_ended JSON messages interleaved with the transcription stream on the same connection. Boundaries are derived from the audio signal and are independent of transcript finalization (is_final); they do not coincide with word boundaries. WebSocket only.

vadanyOptional

Alias for vad_events. When both are set, vad_events takes precedence and vad is ignored.

endpointinganyOptionalDefaults to true

Finalize the current utterance on trailing silence. On by default. With endpointing=true, eou_timeout_ms sets the silence window. Set endpointing=false to finalize on the model's end-of-utterance timer instead.

eou_timeout_msstringOptional

With endpointing=true, the trailing-silence window in milliseconds, 600 if unset. With endpointing=false, the model's end-of-utterance timer, 800 if unset. Range 100 to 10000.

formatanyOptionalDefaults to true

Master formatting switch for transcript responses. When false, forces punctuate=false, capitalize=false, and also disables Inverse Text Normalization (ITN) so it cannot silently reintroduce punctuation or casing.

When true, the punctuate and capitalize params take effect independently. Leave format=true and use those two to fine-tune.

punctuateanyOptionalDefaults to true

When false, strips end-of-sentence punctuation (., ,, ?, !) from the final transcript, words[].word, and utterances[].transcript. Does not affect casing — use capitalize for that.

Overridden to false when format=false.

capitalizeanyOptionalDefaults to true

When false, lowercases the transcript output (final transcript, words[].word, and utterances[].transcript). Some acronyms and proper-noun tokens may be preserved by the model. Does not affect punctuation, use punctuate for that.

Overridden to false when format=false.

itn_normalizeanyOptionalDefaults to false

Enable Inverse Text Normalization to convert spoken-form entities (numbers, dates, currencies, phone numbers, etc.) into written form in finalized transcripts. Example: five five five one two three becomes 5551234.

Overridden to false when format=false.

finalize_on_wordsanyOptionalDefaults to true

Word-count-based finalization.

max_wordsstringOptional
Word count at which the pending transcript is finalized.
redact_piianyOptionalDefaults to false

Redact personally identifiable information from the transcript. Tokens use the shape [ENTITYTYPE_N] where N is a per-entity sequential index (e.g. [FIRSTNAME_1]). Entity types include FIRSTNAME, LASTNAME, PHONENUMBER, ADDRESS, EMAIL.

Language support: effective on language=en and language=hi. Accepted on other codes but redaction is not reliable.

redact_pcianyOptionalDefaults to false

Redact payment card information. Tokens use the shape [ENTITYTYPE_N]. Known entity types include ACCOUNTNAME (cardholder name), CREDITCARDNUMBER (card PAN), CREDITCARDCVV, ZIPCODE, ACCOUNTNUMBER. Use alongside redact_pii=true for combined PII + PCI redaction.

Language support: currently effective only on language=en and language=hi. Setting redact_pci=true on other language codes is accepted but does not reliably redact.

sentence_timestampsanyOptionalDefaults to false

Include sentence-level timing in the transcription event under a new utterances array. Each entry carries text, start, end, and (when diarization is on) speaker. Combine with word_timestamps=true to get both word and sentence boundaries.

keywordsstringOptional

Boost recognition of specific words or phrases for this session. Useful for product names, jargon, proper nouns, and other domain-specific terms the model might otherwise mis-transcribe. Max 100 keywords per session.

Also available on the pre-recorded HTTP endpoint (POST /waves/v1/pulse/get_text) as a query parameter with the same format.

Format: a single comma-separated string (not a JSON array). Each entry is a word or phrase, optionally followed by :INTENSIFIER, a numeric boost multiplier that defaults to 1.0 when omitted.

Example: Blackwell:1,Jensen Huang:2

  • Phrases can include spaces (small language model:2).
  • A phrase cannot contain a comma; commas separate entries.
  • Keyword matching is case-sensitive. Provide the exact casing you want in the output.
  • Duplicates: last-wins. NVIDIA:1,NVIDIA:5 is equivalent to NVIDIA:5.
  • Max 100 keywords per session. Sending more returns 400 with keywords too large (max 100) in errors[].

Intensifier. Default 1, recommended tuning range 1 to 3. Above 10 is not recommended: higher values increase the chance of the model inserting the keyword when it was not spoken. Start every keyword at 1 and only raise it if the word is still being missed.

Wire format: pass as a query-string parameter, URL-encoded. URLSearchParams in JavaScript and urlencode() in Python handle the encoding for you. Pass the raw string, not a JSON-encoded array.

Send

sendAudiostringRequiredformat: "binary"
OR
sendFinalizeobjectRequired
Flush the current audio buffer, run ITN over the accumulated utterance, and emit one `is_final` transcript. The WebSocket stays open and accepts audio for the next user turn. Send this once per turn in any multi-turn flow (voice agents, conversational STT).
OR
sendCloseobjectRequired
Flush any remaining buffered audio, emit the terminal `is_final` + `is_last` transcript, then close the WebSocket. Send this once at the end of the session — end of call, app shutdown, or after the entire buffer of a single-shot transcription is streamed.
OR
sendKeepAliveobjectRequired
Reset the connection inactivity timer without sending audio. The server replies with `{"type":"pong"}`. The frame is discarded before model processing. It does not affect the transcript, contains no audio data, and is not billed. It does not extend the 20-minute model-session idle window. For pauses longer than that, close the session and open a new one when audio resumes.

Receive

receiveTranscriptionobjectRequired
OR
receiveTranscriptionobjectRequired
OR
receiveTranscriptionobjectRequired
OR
receiveTranscriptionobjectRequired
OR
receiveTranscriptionobjectRequired