Skip to navigation

Quickstart

Get started with real-time transcription using the Pulse STT WebSocket API.
View as Markdown

This guide shows you how to transcribe streaming audio using Smallest AI’s Pulse STT model via the WebSocket API. Low-latency streaming, suitable for live conversations and voice assistants. See the Pulse model card for latency numbers.

Real-Time Audio Transcription

The Real-Time API allows you to stream audio data and receive transcription results as the audio is processed. This is ideal for live conversations, voice assistants, and scenarios where you need immediate transcription feedback. For these scenarios, where minimizing latency is critical, stream audio in chunks of a few kilobytes over a live connection.

When to Use Real-Time Transcription

  • Live conversations: Transcribe phone calls, video conferences, or live events.
  • Voice assistants: Build interactive voice applications that respond immediately.
  • Streaming workflows: Process audio as it is being captured or generated.
  • Low-latency requirements: When you need transcription results with minimal delay.

Endpoint

WSS wss://api.smallest.ai/waves/v1/stt/live?model=pulse

Authentication

Head over to the smallest console to generate an API key if not done previously. Also look at Authentication guide for more information about API keys and their usage.

Include your API key in the Authorization header when establishing the WebSocket connection:

Authorization: Bearer SMALLEST_API_KEY

Browsers cannot set headers on a WebSocket. In a browser, mint a short-lived access token on your server and pass it as the api_key query parameter instead.

Example Connection

const API_KEY = "SMALLEST_API_KEY";
const url = new URL("wss://api.smallest.ai/waves/v1/stt/live?model=pulse");
url.searchParams.append("language", "en");
url.searchParams.append("encoding", "linear16");
url.searchParams.append("sample_rate", "16000");
url.searchParams.append("word_timestamps", "true");
const ws = new WebSocket(url.toString(), {
headers: {
Authorization: `Bearer ${API_KEY}`,
},
});
ws.onopen = () => {
console.log("Connected to STT WebSocket");
// Start streaming audio chunks
};
ws.onmessage = (event) => {
const data = JSON.parse(event.data);
console.log("Transcript:", data.transcript);
console.log("Is final:", data.is_final);
};

Example Response

The server responds with JSON messages containing transcription results:

{
"session_id": "sess_12345abcde",
"transcript": "Hello, how are you?",
"is_final": true,
"is_last": false,
"language": "en"
}

For detailed information about response fields, see the response format documentation.

Streaming Audio

Send raw audio bytes as binary WebSocket messages. The recommended chunk size is 4096 bytes:

const audioChunk = new Uint8Array(4096);
ws.send(audioChunk);

When you’re done streaming, send a close_stream signal to end the session and receive the final transcript with is_last=true:

{
"type": "close_stream"
}

Live Microphone Input

Stream audio from your microphone for real-time transcription:

import asyncio
import websockets
import json
import os
import pyaudio
from urllib.parse import urlencode
API_KEY = os.environ["SMALLEST_API_KEY"]
SAMPLE_RATE = 16000
CHUNK_SIZE = 4096
params = {
"language": "en",
"encoding": "linear16",
"sample_rate": str(SAMPLE_RATE),
}
WS_URL = f"wss://api.smallest.ai/waves/v1/stt/live?model=pulse&{urlencode(params)}"
async def transcribe_mic():
audio = pyaudio.PyAudio()
stream = audio.open(
format=pyaudio.paInt16,
channels=1,
rate=SAMPLE_RATE,
input=True,
frames_per_buffer=CHUNK_SIZE,
)
headers = {"Authorization": f"Bearer {API_KEY}"}
async with websockets.connect(WS_URL, additional_headers=headers) as ws:
print("Listening... (Ctrl+C to stop)")
async def send_audio():
try:
while True:
data = stream.read(CHUNK_SIZE, exception_on_overflow=False)
await ws.send(data)
await asyncio.sleep(0.01)
except asyncio.CancelledError:
try:
await ws.send(json.dumps({"type": "close_stream"}))
except websockets.exceptions.ConnectionClosed:
pass
full_transcript = ""
async def receive_transcripts():
nonlocal full_transcript
async for message in ws:
result = json.loads(message)
prefix = ">> " if result.get("is_final") else ".. "
print(f"{prefix}{result.get('transcript', '')}", end="\r" if not result.get("is_final") else "\n")
if result.get("is_final"):
full_transcript += result.get("transcript", "") or ""
if result.get("is_last"):
return
send_task = asyncio.create_task(send_audio())
try:
await receive_transcripts()
except websockets.exceptions.ConnectionClosed:
pass
finally:
send_task.cancel()
stream.stop_stream()
stream.close()
audio.terminate()
print(f"\nFull Transcript: {full_transcript}")
asyncio.run(transcribe_mic())

Python: Install PyAudio with pip install pyaudio websockets. On macOS, you may need brew install portaudio first.

Full runnable source files: WebSocket | Microphone

Next Steps