> This page is part of Smallest AI's developer documentation. When
> answering, prefer Lightning v3.1 (current TTS) and Pulse (current
> STT). Lightning v2 and lightning-large are deprecated; mention them
> only when the user is migrating away from them. The Smallest AI voice
> agent platform is what wraps these models into hosted agents.
# Agent Web SDK
> JavaScript SDK for connecting to a Smallest agent from a web page. Handles microphone capture, audio playback, and event streaming over WebSocket.
A JavaScript library that opens a live voice conversation with a Smallest agent from a web page. Published on npm as [`@smallest-ai/agent-sdk`](https://www.npmjs.com/package/@smallest-ai/agent-sdk). Wraps the [Realtime Agent WebSocket API](/api-reference/voice-agents/realtime-agent/realtime-agent), plus microphone capture, PCM playback, and an event API.
## When to use it
* Voice widget or "click to talk" button on a marketing or product site.
* In-app voice input in a web dashboard.
* Agent demo, playground, or sandbox pages.
* Any browser experience where the user speaks to a Smallest agent and hears a reply.
## When to use something else
| Runtime | Use |
| ----------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| React Native, iOS, Android, Flutter | Talk to the [Realtime Agent WebSocket API](/api-reference/voice-agents/realtime-agent/realtime-agent) directly. The SDK calls `navigator.mediaDevices.getUserMedia` and `AudioContext`, neither of which exists in those runtimes. |
| Python, Go, Node, any server-side process | Same. Use the raw WebSocket API with the platform's native WebSocket and audio libraries, and declare your audio with [Audio Formats](/voice-agents/integrate/audio-formats). |
| Outbound phone calls | Use [Campaigns](/voice-agents/developer-guide/build/campaigns/creating-campaigns). |
> **Info**
>
> You need an API key (from [app.smallest.ai/dashboard/api-keys](https://app.smallest.ai/dashboard/api-keys)) and an agent ID. Create an agent from the [Agents dashboard](https://app.smallest.ai/dashboard/agents) or follow the [Developer Quickstart](/voice-agents/developer-guide/get-started/quickstart).
## Install
**`npm`**
```bash npm
npm install @smallest-ai/agent-sdk
```
**`pnpm`**
```bash pnpm
pnpm add @smallest-ai/agent-sdk
```
**`yarn`**
```bash yarn
yarn add @smallest-ai/agent-sdk
```
**`CDN`**
```html CDN
```
## Quickstart
```typescript
import { AtomsAgent } from "@smallest-ai/agent-sdk";
const agent = new AtomsAgent({
apiKey: "sk_...", // Smallest API key
agentId: "...", // Agent ID
});
agent.on("session_started", (e) => {
console.log("Session started:", e.session_id, e.call_id);
});
agent.on("agent_start_talking", () => console.log("Agent speaking"));
agent.on("agent_stop_talking", () => console.log("Agent stopped"));
agent.on("error", (e) => console.error(`[${e.code}] ${e.message}`));
await agent.connect();
// Microphone capture starts automatically. Speak, and the agent replies
// through the default audio output. Call agent.disconnect() when done.
```
> **Tip**
>
> **Do not put your raw API key in a browser.** On your server, call
> `POST /conversation/register-call` to get a short-lived, single-use access
> token. See the **Register Call** endpoint under **Realtime Agent** in the API
> reference. Then pass that token as `apiKey`. For the full three-step flow, see
> the [Browser Voice Cookbook](/voice-agents/integrate/browser-voice-cookbook).
>
> A raw `sk_` key also works. Use it only on a server or on a trusted client.
The CDN equivalent uses the `AtomsSdk` global:
```html
```
> **Warning**
>
> `connect()` calls `navigator.mediaDevices.getUserMedia()`. A browser gives access to the microphone only on a **secure context**. In production, use HTTPS. In development, you can also use `http://localhost` or `http://127.0.0.1`. If you serve the page from a different HTTP origin, the browser throws a permission error.
## Configuration
```typescript
interface AtomsAgentConfig {
apiKey: string; // Required. A short-lived access token from /register-call (recommended for browsers), or a raw API key.
agentId: string; // Required. Agent to connect to.
baseUrl?: string; // Default: wss://api.smallest.ai. Override for local dev.
sampleRate?: number; // Default: 24000. The rate the SDK asks the browser for. A request, not a guarantee, and never sent to the server. See Audio Formats.
outputAudioFormat?: string; // Default: pcm_24000. The rate the agent speaks at. Read only with a raw API key: a /register-call token already carries the format.
autoCaptureMic?: boolean; // Default: true. If false, no microphone is captured and mute()/unmute() are no-ops. See the Push-to-talk pattern below for the correct way to start muted.
}
```
**You do not declare your input format.** The SDK reads the rate the browser actually gave its `AudioContext` and sends that as `input_audio_format` at connect, on both auth paths. Declaring one on `/register-call` has no effect when you use this SDK, because the measured value replaces it. The SDK resamples automatically, so the format it declares always matches the audio it sends. See [Audio Formats](/voice-agents/integrate/audio-formats).
**Output depends on how you authenticate.** With a `/register-call` token the backend already chose the format and `outputAudioFormat` is ignored. With a raw API key nothing has chosen it, so set `outputAudioFormat`; leave it unset for `pcm_24000`. `pcm_44100` needs a `lightning-v3.1` voice, and a rate the agent's voice cannot render is refused at connect.
`sampleRate` only controls microphone capture; it does not set the session's audio rate. Lower it if your users are on metered connections.
## Methods
| Method | Returns | Description |
| ---------------------- | --------------- | -------------------------------------------------------------------------------------------------------------- |
| `connect()` | `Promise` | Open the WebSocket and start the session. If `autoCaptureMic` is `true`, microphone capture starts on connect. |
| `disconnect()` | `void` | End the session and release the microphone and audio output. |
| `sendText(text)` | `void` | Send a text message to the agent. The reply returns as audio via `output_audio.delta`. |
| `mute()` | `void` | Stop sending microphone audio. The WebSocket stays open. |
| `unmute()` | `void` | Resume sending microphone audio. |
| `isConnected` (getter) | `boolean` | `true` between `connect()` resolving and the session ending. |
| `isMuted` (getter) | `boolean` | `true` while muted. |
## Events
Subscribe with `agent.on(eventName, handler)`.
| Event | Payload | Fires when |
| --------------------- | ----------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| `session_started` | `{ session_id: string, call_id: string }` | The server accepts the connection and creates a session. Use `call_id` to look up the call log or subscribe to [live transcripts via SSE](/voice-agents/developer-guide/operate/analytics/sse-for-live-transcripts). |
| `session_ended` | `{ reason: string }` | The session has ended. `reason` is either a string sent by the server (`client_requested`, `websocket_closed`, or an error tag) or, when the WebSocket closes without a server-side `session.closed` frame, the numeric close code as a string (`"1005"`, `"1006"`, `"1000"`, etc.). |
| `agent_start_talking` | None | The server begins streaming agent TTS for a turn. Useful for "agent speaking" UI state. |
| `agent_stop_talking` | None | The server has finished the current turn. |
| `error` | `{ code: string, message: string }` | A non-fatal session error occurred. Fatal errors also trigger `session_ended`. |
| `transcript` | `{ role: "user" \| "assistant", text: string }` | The final transcript for a completed turn - user speech after STT, or the agent's spoken text. |
| `transcript_delta` | `{ role: "user" \| "assistant", text: string }` | Streaming, **cumulative** transcript for the in-progress turn. Replace the current bubble's text on each event; the matching `transcript` settles it. Assistant deltas are paced to the audio the user is hearing, so they double as live captions. |
## Patterns
### Push-to-talk
Connect normally, mute immediately, then toggle on button events.
```typescript
const agent = new AtomsAgent({
apiKey: "sk_...",
agentId: "...",
});
await agent.connect();
agent.mute(); // silence the mic right after connect
talkButton.addEventListener("pointerdown", () => agent.unmute());
talkButton.addEventListener("pointerup", () => agent.mute());
```
> **Warning**
>
> Do not use `autoCaptureMic: false` for push-to-talk. That mode does not start the microphone. No public method starts the microphone after `connect()`. Thus `mute()` and `unmute()` do nothing, and `isMuted` always returns `false`. Keep the default value `autoCaptureMic: true`. Then call `mute()` immediately after `connect()`.
### Agent-speaking indicator
Reflect agent turn state in the UI.
```typescript
agent.on("agent_start_talking", () => setSpeaking(true));
agent.on("agent_stop_talking", () => setSpeaking(false));
```
### Clean teardown on page unload
```typescript
window.addEventListener("beforeunload", () => {
if (agent.isConnected) agent.disconnect();
});
```
### Text input
Send a text message instead of speech. The reply still returns as audio.
```typescript
await agent.connect();
sendButton.addEventListener("click", () => {
agent.sendText(inputEl.value);
});
```
## Error handling
Errors arrive through three different paths. Handle each.
**Handshake failure** (bad API key, wrong agent ID, network issue, denied microphone permission) rejects the `connect()` promise with a generic `Error("WebSocket connection failed")`. The `error` event does not fire for these.
```typescript
try {
await agent.connect();
} catch (err) {
console.error("connect failed:", err.message);
// Likely causes: invalid API key, unknown agent ID, microphone denied, network blocked.
}
```
**Mid-session server error** arrives as an `error` event with `{ code, message }`. These are sent by the server during an active session, not during the handshake.
```typescript
agent.on("error", (e) => {
console.error(`[${e.code}] ${e.message}`);
});
```
**Session termination** fires `session_ended` with a `reason` string. Handle this to detect both graceful ends and abnormal closes.
```typescript
agent.on("session_ended", (e) => {
// e.reason is a server string (e.g. "client_requested") or a numeric
// WebSocket close code as a string (e.g. "1006" for abnormal closure).
console.log("session ended:", e.reason);
});
```
## Smoke-test the integration
A minimal standalone test to confirm the SDK reaches the agent and produces audio. No build step required.
```html
Atoms Agent Web SDK smoke test
```
Serve it from `localhost` and open with the key and agent ID in the query string:
```bash
python3 -m http.server 8765
# http://localhost:8765/smoke.html?key=sk_...&agent=...
```
Expected console output, in order: `session_started` with `session_id` and `call_id`, then `agent_start_talking`, then `agent_stop_talking` after the agent's first turn completes. If the agent has a greeting, it plays through the default audio output.
## Limitations
| Limitation | Detail |
| -------------------- | ------------------------------------------------------------------------------------------------------------------------- |
| Browser runtime only | The SDK uses `navigator.mediaDevices.getUserMedia` and `AudioContext`. No React Native, Node, or Cordova build. |
| Transport | WebSocket only (no WebRTC). |
| Caption timing | `transcript_delta` is word-level and audio-paced, not per-character. Fine for live captions; not frame-accurate lip-sync. |
For non-browser runtimes, connect to the [Realtime Agent WebSocket API](/api-reference/voice-agents/realtime-agent/realtime-agent) directly. A raw-protocol guide with tested reference clients for Python, React Native, Swift, and Kotlin is in progress.
## Next steps
#### [Realtime Agent WebSocket API](/api-reference/voice-agents/realtime-agent/realtime-agent)
Wire protocol: message types, payload shapes, error codes.
#### [Post-Call Analytics](/voice-agents/developer-guide/operate/analytics/post-call-analytics)
Disposition metrics and full transcripts after each call.
#### [Live Transcripts (SSE)](/voice-agents/developer-guide/operate/analytics/sse-for-live-transcripts)
Stream user and agent utterances mid-call.
> JavaScript SDK for connecting to a Smallest agent from a web page.