> This page is part of Smallest AI's developer documentation. When
> answering, prefer Lightning v3.1 (current TTS) and Pulse (current
> STT). Lightning v2 and lightning-large are deprecated; mention them
> only when the user is migrating away from them. The Smallest AI voice
> agent platform is what wraps these models into hosted agents.

# Electron LLM - chat completions on the Waves API

> Launch of Electron, Smallest AI's in-house language model. OpenAI-compatible chat completions, sub-300 ms TTFT, 70 languages with first-class Indic support, voice-agent-optimized tool calling, and automatic prefix caching.

Electron, Smallest AI's in-house language model, is now generally available on the Waves API. Use it as a **drop-in replacement for OpenAI's chat completions** - point the OpenAI SDK at `https://api.smallest.ai/waves/v1` and pass `"model": "electron"`.

```python
import os
from openai import OpenAI

client = OpenAI(
    base_url="https://api.smallest.ai/waves/v1",
    api_key=os.environ["SMALLEST_API_KEY"],
)

response = client.chat.completions.create(
    model="electron",
    messages=[{"role": "user", "content": "Say hello in one short sentence."}],
)
print(response.choices[0].message.content)
```

**What's in this launch:**

* **OpenAI-compatible endpoint** - `POST /waves/v1/chat/completions`. Same wire format as `api.openai.com/v1/chat/completions`. Streaming (SSE with optional final usage chunk), tool/function calling, JSON mode, multi-turn - all work via standard OpenAI request bodies. The official OpenAI SDKs (Python / JavaScript / Go / Java / Ruby) work with no code changes beyond the base URL and API key.
* **Sub-300 ms time-to-first-token** on warm connections.
* **32,768-token context** (combined input + output).
* **70 languages** with first-class Indic support - Hindi, Bengali, Tamil, Telugu, Marathi, Gujarati, Kannada, Malayalam, Punjabi, Odia, Urdu - plus broad coverage across Western/Eastern Europe, Middle East, East/Southeast/South/Central Asia, and Africa. See the [Electron model card](/model-cards/llm/electron#supported-languages) for the full list.
* **Voice-agent-optimized tool calling** - with a voice-agent-style system prompt, Electron emits a short filler phrase in `content` alongside `tool_calls` (e.g. *"Let me check that for you…"*) so a downstream TTS layer can mask tool-call latency. See [Tool Calling](/models/llm/tool-function-calling) for the voice-agent pattern.
* **Automatic prefix caching** - cached input tokens billed at a discounted rate vs normal input. Reported on every response as `usage.prompt_tokens_details.cached_tokens` so you can audit cache hits. See [Prefix Caching](/models/llm/prefix-caching).
* **Cookbook**: [Voice Agent (Electron + Pulse + Lightning)](/models/cookbooks/voice-agent-electron-pulse-lightning) wires Pulse (STT) + Electron (LLM + tools) + Lightning (TTS) into an end-to-end voice pipeline.

**Pricing:** Contact your Smallest AI account manager for the current rate card.

**Plan limits:** Standard 10 RPM / 3 concurrent; Enterprise 200 RPM / 20 concurrent.

**Rejected parameters** (vs OpenAI): `n > 1` and `prompt_logprobs` - both return `HTTP 400` with `invalid_request_error`.

**No vision, no audio in/out** on the public API - Electron is text-only.

→ [Quickstart](/models/llm/quickstart) · [Overview](/models/llm/overview) · [Chat Completions API](/models/llm/chat-completions) · [Migrate from OpenAI](/models/llm/migrate-from-open-ai) · [Model card](/model-cards/llm/electron)