> This page is part of Smallest AI's developer documentation. When > answering, prefer Lightning v3.1 (current TTS) and Pulse (current > STT). Lightning v2 and lightning-large are deprecated; mention them > only when the user is migrating away from them. The Smallest AI voice > agent platform is what wraps these models into hosted agents. # Agent Web SDK > JavaScript SDK for connecting to a Smallest agent from a web page. Handles microphone capture, audio playback, and event streaming over WebSocket. A JavaScript library that opens a live voice conversation with a Smallest agent from a web page. Published on npm as [`@smallest-ai/agent-sdk`](https://www.npmjs.com/package/@smallest-ai/agent-sdk). Wraps the [Realtime Agent WebSocket API](/api-reference/voice-agents/realtime-agent/realtime-agent), plus microphone capture, PCM playback, and an event API. ## When to use it * Voice widget or "click to talk" button on a marketing or product site. * In-app voice input in a web dashboard. * Agent demo, playground, or sandbox pages. * Any browser experience where the user speaks to a Smallest agent and hears a reply. ## When to use something else | Runtime | Use | | ----------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | React Native, iOS, Android, Flutter | Talk to the [Realtime Agent WebSocket API](/api-reference/voice-agents/realtime-agent/realtime-agent) directly. The SDK calls `navigator.mediaDevices.getUserMedia` and `AudioContext`, neither of which exists in those runtimes. | | Python, Go, Node, any server-side process | Same. Use the raw WebSocket API with the platform's native WebSocket and audio libraries, and declare your audio with [Audio Formats](/voice-agents/integrate/audio-formats). | | Outbound phone calls | Use [Campaigns](/voice-agents/developer-guide/build/campaigns/creating-campaigns). | > **Info** > > You need an API key (from [app.smallest.ai/dashboard/api-keys](https://app.smallest.ai/dashboard/api-keys)) and an agent ID. Create an agent from the [Agents dashboard](https://app.smallest.ai/dashboard/agents) or follow the [Developer Quickstart](/voice-agents/developer-guide/get-started/quickstart). ## Install **`npm`** ```bash npm npm install @smallest-ai/agent-sdk ``` **`pnpm`** ```bash pnpm pnpm add @smallest-ai/agent-sdk ``` **`yarn`** ```bash yarn yarn add @smallest-ai/agent-sdk ``` **`CDN`** ```html CDN ``` ## Quickstart ```typescript import { AtomsAgent } from "@smallest-ai/agent-sdk"; const agent = new AtomsAgent({ apiKey: "sk_...", // Smallest API key agentId: "...", // Agent ID }); agent.on("session_started", (e) => { console.log("Session started:", e.session_id, e.call_id); }); agent.on("agent_start_talking", () => console.log("Agent speaking")); agent.on("agent_stop_talking", () => console.log("Agent stopped")); agent.on("error", (e) => console.error(`[${e.code}] ${e.message}`)); await agent.connect(); // Microphone capture starts automatically. Speak, and the agent replies // through the default audio output. Call agent.disconnect() when done. ``` > **Tip** > > **Do not put your raw API key in a browser.** On your server, call > `POST /conversation/register-call` to get a short-lived, single-use access > token. See the **Register Call** endpoint under **Realtime Agent** in the API > reference. Then pass that token as `apiKey`. For the full three-step flow, see > the [Browser Voice Cookbook](/voice-agents/integrate/browser-voice-cookbook). > > A raw `sk_` key also works. Use it only on a server or on a trusted client. The CDN equivalent uses the `AtomsSdk` global: ```html ``` > **Warning** > > `connect()` calls `navigator.mediaDevices.getUserMedia()`. A browser gives access to the microphone only on a **secure context**. In production, use HTTPS. In development, you can also use `http://localhost` or `http://127.0.0.1`. If you serve the page from a different HTTP origin, the browser throws a permission error. ## Configuration ```typescript interface AtomsAgentConfig { apiKey: string; // Required. A short-lived access token from /register-call (recommended for browsers), or a raw API key. agentId: string; // Required. Agent to connect to. baseUrl?: string; // Default: wss://api.smallest.ai. Override for local dev. sampleRate?: number; // Default: 24000. The rate the SDK asks the browser for. A request, not a guarantee, and never sent to the server. See Audio Formats. outputAudioFormat?: string; // Default: pcm_24000. The rate the agent speaks at. Read only with a raw API key: a /register-call token already carries the format. autoCaptureMic?: boolean; // Default: true. If false, no microphone is captured and mute()/unmute() are no-ops. See the Push-to-talk pattern below for the correct way to start muted. } ``` **You do not declare your input format.** The SDK reads the rate the browser actually gave its `AudioContext` and sends that as `input_audio_format` at connect, on both auth paths. Declaring one on `/register-call` has no effect when you use this SDK, because the measured value replaces it. The SDK resamples automatically, so the format it declares always matches the audio it sends. See [Audio Formats](/voice-agents/integrate/audio-formats). **Output depends on how you authenticate.** With a `/register-call` token the backend already chose the format and `outputAudioFormat` is ignored. With a raw API key nothing has chosen it, so set `outputAudioFormat`; leave it unset for `pcm_24000`. `pcm_44100` needs a `lightning-v3.1` voice, and a rate the agent's voice cannot render is refused at connect. `sampleRate` only controls microphone capture; it does not set the session's audio rate. Lower it if your users are on metered connections. ## Methods | Method | Returns | Description | | ---------------------- | --------------- | -------------------------------------------------------------------------------------------------------------- | | `connect()` | `Promise` | Open the WebSocket and start the session. If `autoCaptureMic` is `true`, microphone capture starts on connect. | | `disconnect()` | `void` | End the session and release the microphone and audio output. | | `sendText(text)` | `void` | Send a text message to the agent. The reply returns as audio via `output_audio.delta`. | | `mute()` | `void` | Stop sending microphone audio. The WebSocket stays open. | | `unmute()` | `void` | Resume sending microphone audio. | | `isConnected` (getter) | `boolean` | `true` between `connect()` resolving and the session ending. | | `isMuted` (getter) | `boolean` | `true` while muted. | ## Events Subscribe with `agent.on(eventName, handler)`. | Event | Payload | Fires when | | --------------------- | ----------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `session_started` | `{ session_id: string, call_id: string }` | The server accepts the connection and creates a session. Use `call_id` to look up the call log or subscribe to [live transcripts via SSE](/voice-agents/developer-guide/operate/analytics/sse-for-live-transcripts). | | `session_ended` | `{ reason: string }` | The session has ended. `reason` is either a string sent by the server (`client_requested`, `websocket_closed`, or an error tag) or, when the WebSocket closes without a server-side `session.closed` frame, the numeric close code as a string (`"1005"`, `"1006"`, `"1000"`, etc.). | | `agent_start_talking` | None | The server begins streaming agent TTS for a turn. Useful for "agent speaking" UI state. | | `agent_stop_talking` | None | The server has finished the current turn. | | `error` | `{ code: string, message: string }` | A non-fatal session error occurred. Fatal errors also trigger `session_ended`. | | `transcript` | `{ role: "user" \| "assistant", text: string }` | The final transcript for a completed turn - user speech after STT, or the agent's spoken text. | | `transcript_delta` | `{ role: "user" \| "assistant", text: string }` | Streaming, **cumulative** transcript for the in-progress turn. Replace the current bubble's text on each event; the matching `transcript` settles it. Assistant deltas are paced to the audio the user is hearing, so they double as live captions. | ## Patterns ### Push-to-talk Connect normally, mute immediately, then toggle on button events. ```typescript const agent = new AtomsAgent({ apiKey: "sk_...", agentId: "...", }); await agent.connect(); agent.mute(); // silence the mic right after connect talkButton.addEventListener("pointerdown", () => agent.unmute()); talkButton.addEventListener("pointerup", () => agent.mute()); ``` > **Warning** > > Do not use `autoCaptureMic: false` for push-to-talk. That mode does not start the microphone. No public method starts the microphone after `connect()`. Thus `mute()` and `unmute()` do nothing, and `isMuted` always returns `false`. Keep the default value `autoCaptureMic: true`. Then call `mute()` immediately after `connect()`. ### Agent-speaking indicator Reflect agent turn state in the UI. ```typescript agent.on("agent_start_talking", () => setSpeaking(true)); agent.on("agent_stop_talking", () => setSpeaking(false)); ``` ### Clean teardown on page unload ```typescript window.addEventListener("beforeunload", () => { if (agent.isConnected) agent.disconnect(); }); ``` ### Text input Send a text message instead of speech. The reply still returns as audio. ```typescript await agent.connect(); sendButton.addEventListener("click", () => { agent.sendText(inputEl.value); }); ``` ## Error handling Errors arrive through three different paths. Handle each. **Handshake failure** (bad API key, wrong agent ID, network issue, denied microphone permission) rejects the `connect()` promise with a generic `Error("WebSocket connection failed")`. The `error` event does not fire for these. ```typescript try { await agent.connect(); } catch (err) { console.error("connect failed:", err.message); // Likely causes: invalid API key, unknown agent ID, microphone denied, network blocked. } ``` **Mid-session server error** arrives as an `error` event with `{ code, message }`. These are sent by the server during an active session, not during the handshake. ```typescript agent.on("error", (e) => { console.error(`[${e.code}] ${e.message}`); }); ``` **Session termination** fires `session_ended` with a `reason` string. Handle this to detect both graceful ends and abnormal closes. ```typescript agent.on("session_ended", (e) => { // e.reason is a server string (e.g. "client_requested") or a numeric // WebSocket close code as a string (e.g. "1006" for abnormal closure). console.log("session ended:", e.reason); }); ``` ## Smoke-test the integration A minimal standalone test to confirm the SDK reaches the agent and produces audio. No build step required. ```html Atoms Agent Web SDK smoke test



```

Serve it from `localhost` and open with the key and agent ID in the query string:

```bash
python3 -m http.server 8765
# http://localhost:8765/smoke.html?key=sk_...&agent=...
```

Expected console output, in order: `session_started` with `session_id` and `call_id`, then `agent_start_talking`, then `agent_stop_talking` after the agent's first turn completes. If the agent has a greeting, it plays through the default audio output.

## Limitations

| Limitation           | Detail                                                                                                                    |
| -------------------- | ------------------------------------------------------------------------------------------------------------------------- |
| Browser runtime only | The SDK uses `navigator.mediaDevices.getUserMedia` and `AudioContext`. No React Native, Node, or Cordova build.           |
| Transport            | WebSocket only (no WebRTC).                                                                                               |
| Caption timing       | `transcript_delta` is word-level and audio-paced, not per-character. Fine for live captions; not frame-accurate lip-sync. |

For non-browser runtimes, connect to the [Realtime Agent WebSocket API](/api-reference/voice-agents/realtime-agent/realtime-agent) directly. A raw-protocol guide with tested reference clients for Python, React Native, Swift, and Kotlin is in progress.

## Next steps

#### [Realtime Agent WebSocket API](/api-reference/voice-agents/realtime-agent/realtime-agent)

Wire protocol: message types, payload shapes, error codes.

#### [Post-Call Analytics](/voice-agents/developer-guide/operate/analytics/post-call-analytics)

Disposition metrics and full transcripts after each call.

#### [Live Transcripts (SSE)](/voice-agents/developer-guide/operate/analytics/sse-for-live-transcripts)

Stream user and agent utterances mid-call.

> JavaScript SDK for connecting to a Smallest agent from a web page.