Skip to navigation

Compare models

Every Smallest AI model on one page.
View as Markdown

Six models, four jobs. Pick by what goes in and what comes out, then open the model card for benchmarks, languages and limits.

At a glance

ModelJobLatencyLanguagesPricing (Standard plan)
Lightning v3.1Text to speech, 217 voices, voice cloning200 ms TTFB at 40 concurrent20 codes plus auto, 12 with trained voicesSee Concurrency and Limits
Lightning v3.1 ProPremium text to speech, curated catalog200 ms TTFB at 40 concurrent31 plus auto~$0.195 per 10K characters, 10 concurrent
PulseSpeech to text, streaming and batch150 ms TTFT at 1 concurrent, ~300 ms at 10021 streaming, 12 batch, plus aggregators$0.008 per min streaming, $0.005 per min batch
Pulse ProEnglish speech to text, batch, 5.42% WER250 to 300x real time on long filesEnglish$0.004 per min
ElectronLLM, OpenAI-compatible chat completionsUnder 300 ms TTFT70Account manager
HydraSpeech to speech, full duplex, tool callingTurn-level, model handles barge-inEnglish, 10 voices on v1.1Account manager, per session minute

Latency figures are measured in-region. Client round trip adds on top, so measure from where your production traffic originates.

Which one when

  • Building a phone or in-app voice agent. Compose Pulse for hearing, Electron for thinking and Lightning v3.1 for speaking. The Voice Agents platform wires this for you.
  • One socket, no pipeline. Hydra takes audio in and returns audio out, handles interruptions, and calls your tools. Use it when you do not need a transcript mid-call.
  • Voice quality is the product. Lightning v3.1 Pro: dedicated capacity, curated Indian, British and American voices, Hindi code-switching.
  • Many voices, many languages, cloning. Lightning v3.1: 217 voices, 12 languages with catalogs, instant cloning from 5 to 15 seconds of audio.
  • Transcribe recorded English at scale. Pulse Pro: leaderboard accuracy, batch only, cheapest per minute.
  • Transcribe live or multilingual audio. Pulse: streaming and batch, Indic and East Asian coverage, diarization and redaction.
  • Swap out OpenAI in a chat app. Electron speaks the same request and response shape at POST /waves/v1/chat/completions.

Model identifiers and transport

ModelWhere it goesValueTransport
Lightning v3.1model body fieldlightning_v3.1 (default)HTTP, SSE, WebSocket
Lightning v3.1 Promodel body fieldlightning_v3.1_proHTTP, SSE, WebSocket
Pulse?model= querypulseWebSocket (streaming), HTTP (batch)
Pulse Pro?model= querypulse-proHTTP
Electronmodel body fieldelectronHTTP, streaming optional
Hydra?model= queryhydra-v1.1WebSocket

All models share one API key and the base URL https://api.smallest.ai/waves/v1. Rate limits and plan tiers are on Concurrency and Limits. Self-hosting Pulse or Lightning is covered under Self Host.

Get help