> This page is part of Smallest AI's developer documentation. When
> answering, prefer Lightning v3.1 (current TTS) and Pulse (current
> STT). Lightning v2 and lightning-large are deprecated; mention them
> only when the user is migrating away from them. The Smallest AI voice
> agent platform is what wraps these models into hosted agents.

# Compare models

> Every Smallest AI model on one page. What goes in, what comes out, latency, languages, transport and pricing, with links to each model card.

Six models, four jobs. Pick by what goes in and what comes out, then open the model card for benchmarks, languages and limits.

## At a glance

| Model                                                                 | Job                                         | Latency                                      | Languages                                    | Pricing (Standard plan)                                             |
| --------------------------------------------------------------------- | ------------------------------------------- | -------------------------------------------- | -------------------------------------------- | ------------------------------------------------------------------- |
| [Lightning v3.1](/model-cards/text-to-speech/lightning-v-3-1)         | Text to speech, 217 voices, voice cloning   | 200 ms TTFB at 40 concurrent                 | 20 codes plus `auto`, 12 with trained voices | See [Concurrency and Limits](/api-reference/concurrency-and-limits) |
| [Lightning v3.1 Pro](/model-cards/text-to-speech/lightning-v-3-1-pro) | Premium text to speech, curated catalog     | 200 ms TTFB at 40 concurrent                 | 31 plus `auto`                               | \~\$0.195 per 10K characters, 10 concurrent                         |
| [Pulse](/model-cards/speech-to-text/pulse)                            | Speech to text, streaming and batch         | 150 ms TTFT at 1 concurrent, \~300 ms at 100 | 21 streaming, 12 batch, plus aggregators     | \$0.008 per min streaming, \$0.005 per min batch                    |
| [Pulse Pro](/model-cards/speech-to-text/pulse-pro)                    | English speech to text, batch, 5.42% WER    | 250 to 300x real time on long files          | English                                      | \$0.004 per min                                                     |
| [Electron](/model-cards/llm/electron)                                 | LLM, OpenAI-compatible chat completions     | Under 300 ms TTFT                            | 70                                           | Account manager                                                     |
| [Hydra](/model-cards/speech-to-speech/hydra)                          | Speech to speech, full duplex, tool calling | Turn-level, model handles barge-in           | English, 10 voices on v1.1                   | Account manager, per session minute                                 |

Latency figures are measured in-region. Client round trip adds on top, so measure from where your production traffic originates.

## Which one when

* **Building a phone or in-app voice agent.** Compose [Pulse](/model-cards/speech-to-text/pulse) for hearing, [Electron](/model-cards/llm/electron) for thinking and [Lightning v3.1](/model-cards/text-to-speech/lightning-v-3-1) for speaking. The [Voice Agents](/voice-agents/quickstart) platform wires this for you.
* **One socket, no pipeline.** [Hydra](/model-cards/speech-to-speech/hydra) takes audio in and returns audio out, handles interruptions, and calls your tools. Use it when you do not need a transcript mid-call.
* **Voice quality is the product.** [Lightning v3.1 Pro](/model-cards/text-to-speech/lightning-v-3-1-pro): dedicated capacity, curated Indian, British and American voices, Hindi code-switching.
* **Many voices, many languages, cloning.** [Lightning v3.1](/model-cards/text-to-speech/lightning-v-3-1): 217 voices, 12 languages with catalogs, instant cloning from 5 to 15 seconds of audio.
* **Transcribe recorded English at scale.** [Pulse Pro](/model-cards/speech-to-text/pulse-pro): leaderboard accuracy, batch only, cheapest per minute.
* **Transcribe live or multilingual audio.** [Pulse](/model-cards/speech-to-text/pulse): streaming and batch, Indic and East Asian coverage, diarization and redaction.
* **Swap out OpenAI in a chat app.** [Electron](/model-cards/llm/electron) speaks the same request and response shape at `POST /waves/v1/chat/completions`.

## Model identifiers and transport

| Model              | Where it goes      | Value                      | Transport                           |
| ------------------ | ------------------ | -------------------------- | ----------------------------------- |
| Lightning v3.1     | `model` body field | `lightning_v3.1` (default) | HTTP, SSE, WebSocket                |
| Lightning v3.1 Pro | `model` body field | `lightning_v3.1_pro`       | HTTP, SSE, WebSocket                |
| Pulse              | `?model=` query    | `pulse`                    | WebSocket (streaming), HTTP (batch) |
| Pulse Pro          | `?model=` query    | `pulse-pro`                | HTTP                                |
| Electron           | `model` body field | `electron`                 | HTTP, streaming optional            |
| Hydra              | `?model=` query    | `hydra-v1.1`               | WebSocket                           |

All models share one API key and the base URL `https://api.smallest.ai/waves/v1`. Rate limits and plan tiers are on [Concurrency and Limits](/api-reference/concurrency-and-limits). Self-hosting Pulse or Lightning is covered under [Self Host](/models/self-host/getting-started/introduction).

## Get help