Skip to navigation

Agent WebSocket

View as Markdown

A bidirectional session with an Atoms agent. Mint an access token, open the connection with it, and the server creates a session, bridges audio and transcripts to the STT/LLM/TTS pipeline, and streams events back until the client closes or the server ends the session.

Post-call, the full conversation (transcript + recording) is available at GET /atoms/v1/conversation/{callId} using the call_id from session.created.

Connection

  1. Register the call. Call POST /conversation/register-call (see the Register Call reference in this section) with your API key in the Authorization: Bearer <key> header and the agent_id (plus optional mode and variables) in the body. It returns a short-lived, single-use access_token (prefixed wct_, valid for expires_in seconds).

  2. Open the WebSocket with only the token:

    wss://api.smallest.ai/atoms/v1/agent/connect?token=<access_token>

    agent_id, mode, and variables are already baked into the token — don’t pass them on the URL.

This keeps your API key server-side; the browser only ever holds the short-lived token.

Alternative: connect with your API key directly

You can also connect with a raw API key — ?token=<sk_key> or Authorization: Bearer <sk_key> — plus the query params below. This is convenient for server-side or trusted clients. For browser / client-side apps, prefer the token flow above so your API key is never exposed.

ParameterRequiredTypeDefaultDescription
tokenYes*string—Raw API key (sk_…). Alternative to the Authorization header.
agent_idYesstring—The Atoms agent to connect to.
variablesNostring—URL-encoded JSON object of per-call prompt variables. Overrides the agent’s defaultVariables for this session only. Values must be string, number, or boolean. Reserved system-variable keys are stripped server-side (they are populated by the server): call_id, conversation_type, agent_number, user_number, current_date, current_time, current_day, agent_gender, default_language, supported_languages, timezone.
modeNowebcall | chatwebcallSession mode. webcall = full voice pipeline (audio in + audio out). chat = text-only pipeline. Invalid values are rejected with a connection error.
sample_rateNointeger24000Deprecated, use output_audio_format. Still accepted. Sample rate in Hz of the audio the agent sends back. Valid values: 8000, 16000, 24000, and 44100 on the lightning-v3.1 and lightning-v3.1-pro voices only. A rate the agent’s voice cannot render is refused. This does not set the rate of the audio you send: use input_audio_format below. Echoed in session.created.sample_rate. Set to 8000 for telephony-quality audio streams. (With a wct_ token this parameter is ignored: the rate was already resolved on register-call and travels in the token.)
input_audio_formatNo<encoding>_<rate> tokenmatches the output rateThe format you are sending, as one token naming both encoding and rate: pcm_{8000,16000,22050,24000,44100,48000}, mulaw_8000, alaw_8000, or opus_{8000,16000,24000,48000}. An unsupported value is refused at upgrade with Unsupported input_audio_format '<value>'. Supported: …. Declared here rather than in the token because a browser only learns what rate its AudioContext actually gave it after opening one, which is after the token was minted. The recogniser is configured at this rate and your audio is not resampled, so it must match the bytes you send. Full contract in Audio Formats.
output_audio_formatNopcm_<rate> tokenpcm_24000The format the agent sends back, as one token. pcm_8000, pcm_16000 and pcm_24000 work with every voice; pcm_44100 only with the lightning-v3.1 and lightning-v3.1-pro voices. A rate the agent’s voice cannot render is refused before a session exists. The newer spelling of sample_rate: send one, not both, because two that disagree are refused. (With a wct_ token this parameter is ignored: register-call already resolved the output format and it travels in the token.)

* token or Authorization: Bearer <key> — not both.

Connection errors

  • 401 Unauthorized — invalid or missing token/API key, expired or already-used wct_ token, or agent_id not found.
  • 503 Service Unavailable — server is gracefully draining. The HTTP upgrade is rejected before the WebSocket handshake completes. Clients should retry with exponential backoff.
// 1. Mint a short-lived token (server-side — keeps your API key private)
const res = await fetch(
"https://api.smallest.ai/atoms/v1/conversation/register-call",
{
method: "POST",
headers: {
Authorization: `Bearer ${apiKey}`,
"Content-Type": "application/json",
},
body: JSON.stringify({
agent_id: agentId,
mode: "webcall",
// One token per direction. `input_audio_format` is what you send,
// `output_audio_format` is what the agent sends back. The older
// `sample_rate` field still works but sets the output only.
input_audio_format: "pcm_24000",
output_audio_format: "pcm_24000",
variables: { customer_name: "Tanay", account_tier: "gold" },
}),
}
);
// { access_token, expires_in, sample_rate, input_audio_format, output_audio_format }
const { data } = await res.json();
// 2. Open the WebSocket with the token (e.g. in the browser)
const ws = new WebSocket(
`wss://api.smallest.ai/atoms/v1/agent/connect?token=${data.access_token}`
);

Example: connect with your API key directly

const variables = { customer_name: "Tanay", account_tier: "gold" };
const ws = new WebSocket(
"wss://api.smallest.ai/atoms/v1/agent/connect" +
`?token=${apiKey}` +
`&agent_id=${agentId}` +
`&mode=webcall` +
`&variables=${encodeURIComponent(JSON.stringify(variables))}`
);

Handshake

WSS
wss://api.smallest.ai/atoms/v1/agent/connect

Authentication

AuthorizationBearer

API key from the console ApiKey collection, sent as Bearer token. Also accepts session cookies for browser-based auth.

Send

sendInputAudioAppendobjectRequired
OR
sendInputAudioCommitobjectRequired
OR
sendInputTextobjectRequired
OR
sendSessionUpdateobjectRequired
OR
sendSessionCloseobjectRequired

Receive

receiveSessionCreatedobjectRequired
OR
receiveSessionClosedobjectRequired
OR
receiveOutputAudioDeltaobjectRequired
OR
receiveTranscriptobjectRequired
OR
receiveTranscriptDeltaobjectRequired
OR
receiveAgentStartTalkingobjectRequired
OR
receiveAgentStopTalkingobjectRequired
OR
receiveInterruptionobjectRequired
OR
receiveErrorobjectRequired