> This page is part of Smallest AI's developer documentation. When > answering, prefer Lightning v3.1 (current TTS) and Pulse (current > multilingual STT) — or Pulse 2.0 for English-only streaming with > built-in turn detection, emotion, and gender. Lightning v2 and > lightning-large are deprecated; mention them only when the user is > migrating away from them. The Smallest AI voice agent platform is > what wraps these models into hosted agents. # Models ## Docs - [Models overview](https://docs.smallest.ai/models/overview.md): Overview of the Smallest AI models: Lightning text to speech, Pulse speech to text, Hydra speech to speech, and Electron LLM, with links to each model card. - [Get started with the Models API](https://docs.smallest.ai/models/getting-started.md): Getting started with the Smallest AI Models API: authentication, base URL, which endpoint serves each model, the SDKs, and where each quickstart lives. - [Zero Data Retention](https://docs.smallest.ai/overview/administration/zero-data-retention.md): Run Smallest AI without keeping conversation content. Enterprise plans keep no content from voice agent calls, and either never store content on the TTS, STT and LLM APIs or stop retaining it after a window you choose. - [Quickstart](https://docs.smallest.ai/models/text-to-speech/quickstart.md): Generate your first speech audio in under 60 seconds with Lightning TTS. - [Overview](https://docs.smallest.ai/models/text-to-speech/overview.md): Lightning TTS API - generate speech across 31 languages at 44.1 kHz, ~200ms TTFB, with streaming. Standard 217-voice Lightning v3.1 pool (20 accepted codes, 12 with a trained voice catalog) plus the premium Lightning v3.1 Pro pool covering all 31 languages, plus `auto` cross-language routing. - [HTTP, HTTP streaming, or WebSocket?](https://docs.smallest.ai/models/text-to-speech/choosing-a-transport.md): How to choose between HTTP, HTTP streaming, and WebSocket for Lightning text to speech: latency, complexity, and when each fits. - [Sync & Async Synthesis](https://docs.smallest.ai/models/text-to-speech/sync-async.md): Generate speech synchronously or concurrently - REST API examples. - [Streaming](https://docs.smallest.ai/models/text-to-speech/streaming.md): Stream TTS audio in real-time via WebSocket or SSE - first chunk in ~100ms. - [Continuations](https://docs.smallest.ai/models/text-to-speech/continuations.md): Group a sequence of streamed text fragments into one continuous synthesis on the Lightning TTS WebSocket, so prosody carries across chunks instead of resetting per-request. - [Word-level timestamps](https://docs.smallest.ai/models/text-to-speech/word-timestamps.md): Per-word timing events interleaved with audio chunks on the Lightning TTS WebSocket - for live captions, karaoke highlighting, and avatar lip-sync. - [Math notation](https://docs.smallest.ai/models/text-to-speech/math-notation.md): Opt-in Lightning TTS flag that reads digit-flanked math operators (x, ×, ÷, spaced -, ^, =) as spoken words instead of leaving them for the number reader. Off by default because dimensions, 24x7, and vehicle-reg codes look like math on paper. - [Content filter](https://docs.smallest.ai/models/text-to-speech/content-filter.md): Opt-in profanity filter for Lightning TTS input. Pass content_filter to reject or flag requests whose text matches a profanity word list. Off by default, and your text is never rewritten. - [Pronunciation Dictionaries](https://docs.smallest.ai/models/text-to-speech/pronunciation-dictionaries.md): Learn how to create and use pronunciation dictionaries to control how specific words are pronounced in your text-to-speech synthesis - [Voices, Voice IDs, and Supported Languages](https://docs.smallest.ai/models/text-to-speech/voices-languages.md): Find your voice ID, list available voices, filter by language and accent, and find the right voice for your use case. - [TTS Best Practices](https://docs.smallest.ai/models/text-to-speech/best-practices.md): Voice-agent prompting patterns and text formatting rules that make AI agents sound natural through a text-to-speech engine. - [Lightning benchmarks](https://docs.smallest.ai/models/text-to-speech/benchmarks/performance.md): Overview of Lightning v3.1 and Lightning v3.1 Pro text-to-speech benchmarks: latency, listener ratings, pronunciation accuracy and MOS. - [Latency](https://docs.smallest.ai/models/text-to-speech/benchmarks/latency.md): Lightning v3.1 and Lightning v3.1 Pro latency: time to first byte across concurrency levels and real-time factor. - [Listener ratings](https://docs.smallest.ai/models/text-to-speech/benchmarks/listener-ratings.md): Lightning v3.1 and Lightning v3.1 Pro head-to-head listener ratings for naturalness, expressiveness and delivery against other text-to-speech providers. - [Accuracy](https://docs.smallest.ai/models/text-to-speech/benchmarks/accuracy.md): Lightning text-to-speech pronunciation accuracy measured by transcribing the output with Whisper: jiwer word error rate and an LLM-judged score. - [MOS](https://docs.smallest.ai/models/text-to-speech/benchmarks/mos.md): Lightning v3.1 and Lightning v3.1 Pro mean opinion score (MOS v2) compared with other text-to-speech providers. - [Metrics Overview](https://docs.smallest.ai/models/text-to-speech/benchmarks/metrics-overview.md): Definitions of every Lightning v3.1 TTS quality and latency metric - Naturalness, Expressiveness, Delivery, Accuracy, and MOS variants. - [Quickstart](https://docs.smallest.ai/models/speech-to-text/quickstart.md): Transcribe your first audio file in under 60 seconds with Pulse Pro. - [Overview](https://docs.smallest.ai/models/speech-to-text/overview.md): Smallest Speech-to-Text API. Pulse, Pulse 2.0, and Pulse Pro models behind one unified endpoint with multilingual streaming, built-in turn detection, leaderboard-ranked English accuracy, diarization, word timestamps, and emotion detection. - [Quickstart](https://docs.smallest.ai/models/speech-to-text/pre-recorded/quickstart.md): Transcribe pre-recorded audio files using the unified STT endpoint with Pulse or Pulse Pro - [Audio Specifications](https://docs.smallest.ai/models/speech-to-text/pre-recorded/audio-formats.md): Supported formats, codecs, and recommendations for pre-recorded audio - [Webhooks](https://docs.smallest.ai/models/speech-to-text/pre-recorded/webhooks.md): Receive asynchronous Pulse STT results without polling - [Features](https://docs.smallest.ai/models/speech-to-text/pre-recorded/features.md): Available features for Pre-Recorded Pulse STT API - [Troubleshooting](https://docs.smallest.ai/models/speech-to-text/pre-recorded/troubleshooting.md): Resolve common issues when uploading pre-recorded audio to Pulse STT - [Best Practices](https://docs.smallest.ai/models/speech-to-text/pre-recorded/best-practices.md): Prepare audio inputs before submitting them to Pulse STT - [Code Examples](https://docs.smallest.ai/models/speech-to-text/pre-recorded/code-examples.md): Complete Python examples for transcribing pre-recorded audio with Pulse Pro and Pulse - [Quickstart](https://docs.smallest.ai/models/speech-to-text/realtime-web-socket/quickstart.md): Get started with real-time transcription using the Pulse STT WebSocket API - [Response Format](https://docs.smallest.ai/models/speech-to-text/realtime-web-socket/response-format.md): Understanding the structure and fields of real-time transcription responses - [Audio Specifications](https://docs.smallest.ai/models/speech-to-text/realtime-web-socket/audio-formats.md): Supported audio encoding formats and requirements for real-time WebSocket transcription - [Features](https://docs.smallest.ai/models/speech-to-text/realtime-web-socket/features.md): Available features for Real-Time Pulse STT WebSocket API - [Troubleshooting](https://docs.smallest.ai/models/speech-to-text/realtime-web-socket/troubleshooting.md): Common issues and solutions for real-time WebSocket transcription - [Best Practices](https://docs.smallest.ai/models/speech-to-text/realtime-web-socket/best-practices.md): Optimize your real-time WebSocket transcription for low latency and high accuracy - [Code Examples](https://docs.smallest.ai/models/speech-to-text/realtime-web-socket/code-examples.md): Complete code examples for real-time WebSocket transcription in Python, Node.js, and Browser JavaScript - [Word timestamps](https://docs.smallest.ai/models/speech-to-text/features/word-timestamps.md): Return word-level timing metadata from Pulse STT - [Language detection](https://docs.smallest.ai/models/speech-to-text/features/language-detection.md): Automatically detect and transcribe across the European, North Indian, South Indian, or East Asian Pulse STT language sets - [Sentence-level timestamps](https://docs.smallest.ai/models/speech-to-text/features/utterances.md): Use the utterances array to capture longer segments with speaker labels - [Speaker diarization](https://docs.smallest.ai/models/speech-to-text/features/diarization.md): Attach zero-indexed speaker labels to each word and utterance in a transcript on Pulse batch and streaming. - [PII and PCI Redaction](https://docs.smallest.ai/models/speech-to-text/features/redaction.md): Automatically redact sensitive information from transcriptions - [Gender detection](https://docs.smallest.ai/models/speech-to-text/features/gender-detection.md): Predict speaker gender alongside every transcription - [Emotion detection](https://docs.smallest.ai/models/speech-to-text/features/emotion-detection.md): Capture per-emotion confidence scores from Pulse STT responses - [Keyword Boosting](https://docs.smallest.ai/models/speech-to-text/features/keyword-boosting.md): Boost specific words or phrases so the Pulse speech-to-text model recognizes them correctly, on both streaming and pre-recorded transcription. - [Punctuation Formatting](https://docs.smallest.ai/models/speech-to-text/features/punctuation-formatting.md): Control punctuation and capitalization formatting in real-time transcripts - [End-of-Utterance Timeout](https://docs.smallest.ai/models/speech-to-text/features/end-of-utterance-timeout.md): Control how long Pulse waits after speech ends before finalizing the transcript - [Inverse Text Normalization (ITN)](https://docs.smallest.ai/models/speech-to-text/features/inverse-text-normalization.md): Convert spoken-form transcripts into written form in real time - [Finalize Control](https://docs.smallest.ai/models/speech-to-text/features/finalize-control.md): Word-count finalization parameters on the Pulse STT WebSocket: `finalize_on_words` and `max_words`. - [Voice Activity (VAD) events](https://docs.smallest.ai/models/speech-to-text/features/vad-events.md): Acoustic `speech_started` / `speech_ended` events emitted on the Pulse STT WebSocket alongside `transcription` messages when `vad_events=true` is set on the connection. - [Finalization and Endpointing](https://docs.smallest.ai/models/speech-to-text/features/endpointing.md): How the Pulse STT WebSocket decides when a turn is complete. Covers endpointing, eou_timeout_ms, finalize_on_words, and the client-side finalize / close_stream signals in one place. - [Keep-Alive](https://docs.smallest.ai/models/speech-to-text/features/keep-alive.md): Hold an idle Pulse STT WebSocket open with a `ping` control frame, and understand the inactivity, session, and lifetime limits that apply when no audio is flowing. - [Pulse benchmarks](https://docs.smallest.ai/models/speech-to-text/benchmarks/performance.md): Overview of Pulse, Pulse 2.0, and Pulse Pro speech-to-text benchmarks: Open ASR Leaderboard, latency, FLEURS, ESB, Hindi, East Asian, WildASR, contact center, perturbation, diarization, end-of-turn detection, emotion, and gender. - [Open ASR Leaderboard](https://docs.smallest.ai/models/speech-to-text/benchmarks/open-asr-leaderboard.md): Pulse Pro on the public Open ASR Leaderboard: tied for second at 5.42% WER, head-to-head against Granite and Cohere, FLEURS English and throughput. - [Latency](https://docs.smallest.ai/models/speech-to-text/benchmarks/latency.md): Pulse streaming latency: time to first transcript at 1 to 100 concurrent sessions, measured in-region. - [FLEURS](https://docs.smallest.ai/models/speech-to-text/benchmarks/fleurs.md): Pulse accuracy on the FLEURS multilingual benchmark: streaming English, pre-recorded and streaming per-language word error rates. - [ESB English](https://docs.smallest.ai/models/speech-to-text/benchmarks/esb-english.md): Pulse streaming accuracy on the ESB English benchmark suite: AMI, Earnings22, GigaSpeech, LibriSpeech, SPGISpeech, TED-LIUM, VoxPopuli. - [Hindi](https://docs.smallest.ai/models/speech-to-text/benchmarks/hindi.md): Pulse streaming accuracy on Hindi across multiple public datasets, compared with other speech-to-text APIs. - [East Asian languages](https://docs.smallest.ai/models/speech-to-text/benchmarks/east-asian.md): Pulse streaming accuracy on Mandarin, Cantonese, Japanese and Korean across public datasets, served from the US region. - [WildASR robustness](https://docs.smallest.ai/models/speech-to-text/benchmarks/wild-asr.md): Pulse streaming robustness on the WildASR dataset: noisy, accented, real-world audio compared with other APIs. - [Contact center calls](https://docs.smallest.ai/models/speech-to-text/benchmarks/contact-center.md): Pulse accuracy on real-world English and Hindi contact-center calls, compared with other speech-to-text APIs. - [Perturbation robustness](https://docs.smallest.ai/models/speech-to-text/benchmarks/perturbation.md): Pulse accuracy under the internal English and Hindi perturbation benchmarks: background noise, speed and pitch shifts, telephony codecs. - [Diarization](https://docs.smallest.ai/models/speech-to-text/benchmarks/diarization.md): Pulse speaker diarization: streaming DER against other speech-to-text APIs, DER on Indic languages, and real-time factor. - [End-of-Turn Detection](https://docs.smallest.ai/models/speech-to-text/benchmarks/end-of-turn-detection.md): Pulse 2.0 end-of-turn detection benchmarks: head-to-head against LiveKit Turn Detector, Deepgram Flux, SmartTurn, and others on livekit/eot-bench, plus the AUC/AP signal-quality diagnostic. - [Emotion Detection](https://docs.smallest.ai/models/speech-to-text/benchmarks/emotion-detection.md): Pulse 2.0 emotion detection benchmarks: held-out corpora generalization F1, nine evaluated emotion classes, and the cost of running through the streaming server. - [Gender Detection](https://docs.smallest.ai/models/speech-to-text/benchmarks/gender-detection.md): Pulse 2.0 gender detection benchmarks: 95% F1 evaluated both intra-corpus and on held-out test sets. - [Metrics Overview](https://docs.smallest.ai/models/speech-to-text/benchmarks/metrics-overview.md): Key Pulse STT metrics for quality and latency. - [Evaluation Walkthrough](https://docs.smallest.ai/models/speech-to-text/benchmarks/evaluation-walkthrough.md): Step-by-step guide to evaluate Pulse STT accuracy and performance against your own dataset using the Pulse REST API. - [Measuring Pulse Latency](https://docs.smallest.ai/models/speech-to-text/benchmarks/measuring-latency.md): Measure Pulse streaming + pre-recorded latency using the same metrics the voice-agent industry uses - Time-to-First-Partial, End-of-Turn (EOT), and RTFx - with attribution to the four components of the pipeline. - [Hydra - Speech to Speech](https://docs.smallest.ai/models/speech-to-speech/overview.md): Hydra is Smallest AI's realtime, full-duplex speech-to-speech model. Audio in, audio out, over a single WebSocket - built for phone-grade voice agents. - [Quickstart](https://docs.smallest.ai/models/speech-to-speech/quickstart.md): Talk to Hydra in under a minute - a one-clone browser client with multi-agent presets, tool execution, and a live wire log. - [WebSocket connection](https://docs.smallest.ai/models/speech-to-speech/web-socket-connection.md): Connect to the Hydra speech-to-speech WebSocket - auth, query parameters, idle timeout, close codes. Python, Node, and browser snippets. - [Managing sessions](https://docs.smallest.ai/models/speech-to-speech/managing-sessions.md): Hydra session lifecycle - session.configure, voices, persona, generate_initial_response, mid-session updates, and conversation items. - [Audio I/O](https://docs.smallest.ai/models/speech-to-speech/audio-i-o.md): Audio formats for Hydra - input PCM16 16 kHz, output rate negotiation, chunk sizing, and AudioWorklet patterns for browser clients. - [Turn detection & barge-in](https://docs.smallest.ai/models/speech-to-speech/turn-detection-barge-in.md): How Hydra detects user speech, the events it emits per turn, and how to handle barge-in cleanly on the client. - [Tool calling](https://docs.smallest.ai/models/speech-to-speech/tool-calling.md): Declare tools in session.configure, stream arguments via response.function_call_arguments, post results back, and let Hydra narrate the answer. - [Prompting voice agents](https://docs.smallest.ai/models/speech-to-speech/prompting-voice-agents.md): How to write Hydra system prompts that produce natural-sounding voice agents - persona, length discipline, tool-call prompting, and turn-taking. - [Errors & reconnection](https://docs.smallest.ai/models/speech-to-speech/errors-reconnection.md): Hydra error frame structure, code reference, and reconnection strategy. Errors are diagnostic - the session stays usable unless a close follows. - [Performance](https://docs.smallest.ai/models/speech-to-speech/benchmarks/performance.md): Hydra head-to-head against eight production realtime voice and speech-to-speech models on the AIEWF S2S benchmark - pass rate and voice-to-voice latency, with and without tool calls. - [Metrics Overview](https://docs.smallest.ai/models/speech-to-speech/benchmarks/metrics-overview.md): Definitions of the metrics Hydra is benchmarked on - pass rate, voice-to-voice latency (tool and non-tool), and how each is computed. - [Quickstart](https://docs.smallest.ai/models/llm/quickstart.md): Make your first Electron chat completion in under 60 seconds - point the OpenAI SDK at Smallest AI and go. - [Overview](https://docs.smallest.ai/models/llm/overview.md): Electron - Smallest AI's in-house language model. OpenAI-compatible chat completions, 70 languages with first-class Indic support, voice-agent-optimized tool calling, prefix caching. - [Chat Completions](https://docs.smallest.ai/models/llm/chat-completions.md): POST /waves/v1/chat/completions - OpenAI-compatible chat completion API. Full request and response reference. - [Streaming](https://docs.smallest.ai/models/llm/streaming.md): Stream chat completion tokens via Server-Sent Events - the same wire format as OpenAI. Includes optional usage block for accurate billing on disconnects. - [Tool / Function Calling](https://docs.smallest.ai/models/llm/tool-function-calling.md): Standard OpenAI tools API on Electron, with voice-agent-optimized filler-phrase behavior. Reduces perceived latency on tool calls in conversational pipelines. - [Prefix Caching](https://docs.smallest.ai/models/llm/prefix-caching.md): Cached input tokens are billed at a discounted rate. Structure your prompts to reuse stable prefixes (system prompts, RAG context, conversation history) and pay less. - [Supported Parameters](https://docs.smallest.ai/models/llm/supported-parameters.md): What flows through to Electron, what's rejected, and how Electron extends or restricts the OpenAI Chat Completions request body. - [Migrate from OpenAI](https://docs.smallest.ai/models/llm/migrate-from-open-ai.md): Drop-in replacement for OpenAI's Chat Completions API. Two strings change: the base URL and the API key. - [Best Practices](https://docs.smallest.ai/models/llm/best-practices.md): How to get the best out of Electron - prompt structure, caching, tool calling, streaming patterns, error handling, and cost control. - [Instant Voice Clone (Web UI)](https://docs.smallest.ai/models/voice-cloning/instant-clone-ui.md): Clone a voice from a short audio sample using the Smallest AI console. - [Instant Voice Clone (REST API)](https://docs.smallest.ai/models/voice-cloning/instant-clone-api.md): Create a voice clone in a single HTTP call using the unified voice cloning API. - [Instant Voice Clone (Python SDK)](https://docs.smallest.ai/models/voice-cloning/instant-clone-python-sdk.md): Clone a voice from a short audio sample using the Python SDK. - [Voice Cloning Best Practices](https://docs.smallest.ai/models/voice-cloning/best-practices.md): Guidelines for recording reference audio and achieving high-quality voice clones. - [Speech to Text Examples](https://docs.smallest.ai/models/cookbooks/speech-to-text.md): Production-ready code examples for Pulse STT - from real-time streaming to batch transcription. - [Text to Speech Examples](https://docs.smallest.ai/models/cookbooks/text-to-speech.md): Production-ready code examples for Lightning TTS - from basic synthesis to streaming, voice cloning, and full applications. - [Voice Agent (Electron + Pulse + Lightning)](https://docs.smallest.ai/models/cookbooks/voice-agent-electron-pulse-lightning.md): Build an end-to-end voice agent on Smallest AI: Pulse for transcription, Electron for the LLM brain (with tool calling), Lightning for speech output. - [Error reference](https://docs.smallest.ai/models/troubleshooting/error-reference.md): HTTP status codes returned by Waves TTS and Pulse STT endpoints, with fixes for each. - [Introduction](https://docs.smallest.ai/models/self-host/getting-started/introduction.md): Deploy high-performance speech-to-text and text-to-speech models in your own infrastructure - [Prerequisites](https://docs.smallest.ai/models/self-host/getting-started/prerequisites.md): What you need before deploying Smallest Self-Host - [Why Self-Host?](https://docs.smallest.ai/models/self-host/getting-started/why-self-host.md): Understand when self-hosting our models makes sense for your organization - [Architecture Overview](https://docs.smallest.ai/models/self-host/getting-started/architecture.md): Understanding the components and architecture of Smallest Self-Host deployments - [Hardware Requirements](https://docs.smallest.ai/models/self-host/docker-setup/stt-deployment/prerequisites/hardware-requirements.md): Hardware specifications for deploying Speech-to-Text with Docker - [Software Requirements](https://docs.smallest.ai/models/self-host/docker-setup/stt-deployment/prerequisites/software-requirements.md): Software and dependencies for deploying Speech-to-Text with Docker - [Credentials & Access](https://docs.smallest.ai/models/self-host/docker-setup/stt-deployment/prerequisites/credentials.md): License keys and registry credentials for STT Docker deployment - [Verification Checklist](https://docs.smallest.ai/models/self-host/docker-setup/stt-deployment/prerequisites/verification.md): Verify all prerequisites before deploying STT with Docker - [Quick Start](https://docs.smallest.ai/models/self-host/docker-setup/stt-deployment/quick-start.md): Deploy Smallest Self-Host Speech-to-Text with Docker Compose in under 15 minutes - [Cloud Deployment](https://docs.smallest.ai/models/self-host/docker-setup/stt-deployment/cloud-deployment.md): Recommended GPU instance types on AWS, GCP, and Azure for self-hosting Pulse and Pulse Pro. - [Parallelism and Latency](https://docs.smallest.ai/models/self-host/docker-setup/stt-deployment/parallelism-and-latency.md): Throughput, real-time factor (RTFx), and latency figures for self-hosted Pulse and Pulse Pro deployments. - [Services Overview](https://docs.smallest.ai/models/self-host/docker-setup/stt-deployment/services-overview.md): Detailed breakdown of each service component in the STT Docker deployment - [Configuration](https://docs.smallest.ai/models/self-host/docker-setup/stt-deployment/configuration.md): Advanced configuration options for STT Docker deployments - [Multi-checkpoint deployment](https://docs.smallest.ai/models/self-host/docker-setup/stt-deployment/multi-checkpoint-deployment.md): Deploy Pulse STT on-prem with multiple language checkpoints by pairing one API server with one model server, 1:1. Each deployment scales independently. - [Docker Troubleshooting](https://docs.smallest.ai/models/self-host/docker-setup/stt-deployment/troubleshooting.md): Debug common issues and optimize your STT Docker deployment - [Hardware Requirements](https://docs.smallest.ai/models/self-host/docker-setup/tts-deployment/prerequisites/hardware-requirements.md): Hardware specifications for deploying Text-to-Speech with Docker - [Software Requirements](https://docs.smallest.ai/models/self-host/docker-setup/tts-deployment/prerequisites/software-requirements.md): Software and dependencies for deploying Text-to-Speech with Docker - [Credentials & Access](https://docs.smallest.ai/models/self-host/docker-setup/tts-deployment/prerequisites/credentials.md): License keys and registry credentials for TTS Docker deployment - [Verification Checklist](https://docs.smallest.ai/models/self-host/docker-setup/tts-deployment/prerequisites/verification.md): Verify all prerequisites before deploying TTS with Docker - [Quick Start](https://docs.smallest.ai/models/self-host/docker-setup/tts-deployment/quick-start.md): Deploy Smallest Self-Host Text-to-Speech with Docker Compose in under 15 minutes - [Services Overview](https://docs.smallest.ai/models/self-host/docker-setup/tts-deployment/services-overview.md): Detailed breakdown of each service component in the TTS Docker deployment - [Configuration](https://docs.smallest.ai/models/self-host/docker-setup/tts-deployment/configuration.md): Advanced configuration options for TTS Docker deployments - [Docker Troubleshooting](https://docs.smallest.ai/models/self-host/docker-setup/tts-deployment/troubleshooting.md): Debug common issues and optimize your TTS Docker deployment - [Hardware Requirements](https://docs.smallest.ai/models/self-host/kubernetes-setup/prerequisites/hardware-requirements.md): Cluster and hardware specifications for Kubernetes STT deployment - [Software Requirements](https://docs.smallest.ai/models/self-host/kubernetes-setup/prerequisites/software-requirements.md): Tools and software for Kubernetes STT deployment - [Credentials & Access](https://docs.smallest.ai/models/self-host/kubernetes-setup/prerequisites/credentials.md): License keys and registry credentials for Kubernetes STT deployment - [Verification Checklist](https://docs.smallest.ai/models/self-host/kubernetes-setup/prerequisites/verification.md): Verify all prerequisites before deploying STT on Kubernetes - [Quick Start](https://docs.smallest.ai/models/self-host/kubernetes-setup/quick-start.md): Deploy Smallest Self-Host on Kubernetes with Helm - [AWS EKS Setup](https://docs.smallest.ai/models/self-host/kubernetes-setup/aws/eks-setup.md): Create and configure an EKS cluster for Smallest Self-Host with GPU support - [GPU Nodes Configuration](https://docs.smallest.ai/models/self-host/kubernetes-setup/aws/gpu-nodes.md): Advanced GPU node setup and optimization for AWS EKS - [IAM & IRSA](https://docs.smallest.ai/models/self-host/kubernetes-setup/aws/iam-irsa.md): Configure IAM Roles for Service Accounts in EKS - [EFS Configuration](https://docs.smallest.ai/models/self-host/kubernetes-setup/storage-pvc/efs-configuration.md): Set up Amazon EFS for shared storage in AWS EKS - [Model Storage](https://docs.smallest.ai/models/self-host/kubernetes-setup/storage-pvc/model-storage.md): Optimize model storage and caching strategies for Lightning ASR - [Redis Persistence](https://docs.smallest.ai/models/self-host/kubernetes-setup/storage-pvc/redis-persistence.md): Configure Redis data persistence and high availability - [HPA Configuration](https://docs.smallest.ai/models/self-host/kubernetes-setup/autoscaling/hpa-configuration.md): Configure Horizontal Pod Autoscaling based on custom metrics - [Cluster Autoscaler](https://docs.smallest.ai/models/self-host/kubernetes-setup/autoscaling/cluster-autoscaler.md): Automatically scale EKS cluster nodes based on pod resource requirements - [Metrics Setup](https://docs.smallest.ai/models/self-host/kubernetes-setup/autoscaling/metrics-setup.md): Configure Prometheus, ServiceMonitor, and custom metrics collection for Lightning ASR - [Grafana Dashboards](https://docs.smallest.ai/models/self-host/kubernetes-setup/autoscaling/grafana-dashboards.md): Visualize metrics, autoscaling behavior, and system performance - [Kubernetes Troubleshooting](https://docs.smallest.ai/models/self-host/kubernetes-setup/troubleshooting.md): Debug common issues in Kubernetes deployments - [Common Issues](https://docs.smallest.ai/models/self-host/troubleshooting/common-issues.md): Quick solutions to frequently encountered problems - [Debugging Guide](https://docs.smallest.ai/models/self-host/troubleshooting/debugging-guide.md): Advanced debugging techniques for Smallest Self-Host - [Logs Analysis](https://docs.smallest.ai/models/self-host/troubleshooting/logs-analysis.md): Interpret logs and error messages from Smallest Self-Host - [Authentication](https://docs.smallest.ai/models/self-host/api-reference/authentication.md): Authenticate API requests with your license key - [Health Check](https://docs.smallest.ai/models/self-host/api-reference/endpoints/health-check.md): Monitor service health and availability - [Transcription](https://docs.smallest.ai/models/self-host/api-reference/endpoints/transcription.md): Convert speech to text with the /v1/listen endpoint - [Integration Examples](https://docs.smallest.ai/models/self-host/api-reference/examples.md): Complete examples for integrating with Smallest Self-Host