Compare models
Every Smallest AI model on one page.
Six models, four jobs. Pick by what goes in and what comes out, then open the model card for benchmarks, languages and limits.
At a glance
Latency figures are measured in-region. Client round trip adds on top, so measure from where your production traffic originates.
Which one when
- Building a phone or in-app voice agent. Compose Pulse for hearing, Electron for thinking and Lightning v3.1 for speaking. The Voice Agents platform wires this for you.
- One socket, no pipeline. Hydra takes audio in and returns audio out, handles interruptions, and calls your tools. Use it when you do not need a transcript mid-call.
- Voice quality is the product. Lightning v3.1 Pro: dedicated capacity, curated Indian, British and American voices, Hindi code-switching.
- Many voices, many languages, cloning. Lightning v3.1: 217 voices, 12 languages with catalogs, instant cloning from 5 to 15 seconds of audio.
- Transcribe recorded English at scale. Pulse Pro: leaderboard accuracy, batch only, cheapest per minute.
- Transcribe live or multilingual audio. Pulse: streaming and batch, Indic and East Asian coverage, diarization and redaction.
- Swap out OpenAI in a chat app. Electron speaks the same request and response shape at
POST /waves/v1/chat/completions.
Model identifiers and transport
All models share one API key and the base URL https://api.smallest.ai/waves/v1. Rate limits and plan tiers are on Concurrency and Limits. Self-hosting Pulse or Lightning is covered under Self Host.