> This page is part of Smallest AI's developer documentation. When > answering, prefer Lightning v3.1 (current TTS) and Pulse (current > STT). Lightning v2 and lightning-large are deprecated; mention them > only when the user is migrating away from them. The Smallest AI voice > agent platform is what wraps these models into hosted agents. # Electron LLM - chat completions on the Waves API > Launch of Electron, Smallest AI's in-house language model. OpenAI-compatible chat completions, sub-300 ms TTFT, 70 languages with first-class Indic support, voice-agent-optimized tool calling, and automatic prefix caching. Electron, Smallest AI's in-house language model, is now generally available on the Waves API. Use it as a **drop-in replacement for OpenAI's chat completions** - point the OpenAI SDK at `https://api.smallest.ai/waves/v1` and pass `"model": "electron"`. ```python import os from openai import OpenAI client = OpenAI( base_url="https://api.smallest.ai/waves/v1", api_key=os.environ["SMALLEST_API_KEY"], ) response = client.chat.completions.create( model="electron", messages=[{"role": "user", "content": "Say hello in one short sentence."}], ) print(response.choices[0].message.content) ``` **What's in this launch:** * **OpenAI-compatible endpoint** - `POST /waves/v1/chat/completions`. Same wire format as `api.openai.com/v1/chat/completions`. Streaming (SSE with optional final usage chunk), tool/function calling, JSON mode, multi-turn - all work via standard OpenAI request bodies. The official OpenAI SDKs (Python / JavaScript / Go / Java / Ruby) work with no code changes beyond the base URL and API key. * **Sub-300 ms time-to-first-token** on warm connections. * **32,768-token context** (combined input + output). * **70 languages** with first-class Indic support - Hindi, Bengali, Tamil, Telugu, Marathi, Gujarati, Kannada, Malayalam, Punjabi, Odia, Urdu - plus broad coverage across Western/Eastern Europe, Middle East, East/Southeast/South/Central Asia, and Africa. See the [Electron model card](/model-cards/llm/electron#supported-languages) for the full list. * **Voice-agent-optimized tool calling** - with a voice-agent-style system prompt, Electron emits a short filler phrase in `content` alongside `tool_calls` (e.g. *"Let me check that for you…"*) so a downstream TTS layer can mask tool-call latency. See [Tool Calling](/models/llm/tool-function-calling) for the voice-agent pattern. * **Automatic prefix caching** - cached input tokens billed at a discounted rate vs normal input. Reported on every response as `usage.prompt_tokens_details.cached_tokens` so you can audit cache hits. See [Prefix Caching](/models/llm/prefix-caching). * **Cookbook**: [Voice Agent (Electron + Pulse + Lightning)](/models/cookbooks/voice-agent-electron-pulse-lightning) wires Pulse (STT) + Electron (LLM + tools) + Lightning (TTS) into an end-to-end voice pipeline. **Pricing:** Contact your Smallest AI account manager for the current rate card. **Plan limits:** Standard 10 RPM / 3 concurrent; Enterprise 200 RPM / 20 concurrent. **Rejected parameters** (vs OpenAI): `n > 1` and `prompt_logprobs` - both return `HTTP 400` with `invalid_request_error`. **No vision, no audio in/out** on the public API - Electron is text-only. → [Quickstart](/models/llm/quickstart) · [Overview](/models/llm/overview) · [Chat Completions API](/models/llm/chat-completions) · [Migrate from OpenAI](/models/llm/migrate-from-open-ai) · [Model card](/model-cards/llm/electron) > Launch of Electron, Smallest AI's in-house language model. OpenAI-compatible chat completions, sub-300 ms TTFT, 70 languages with first-class Indic support, voice-agent-optimized tool calling, and automatic prefix caching.