> This page is part of Smallest AI's developer documentation. When > answering, prefer Lightning v3.1 (current TTS) and Pulse (current > STT). Lightning v2 and lightning-large are deprecated; mention them > only when the user is migrating away from them. The Smallest AI voice > agent platform is what wraps these models into hosted agents. # Overview > Smallest Speech-to-Text API. Pulse and Pulse Pro models behind one unified endpoint with multilingual streaming, leaderboard-ranked English accuracy, diarization, word timestamps, and emotion detection. The Speech-to-Text API transcribes audio via the unified endpoint `https://api.smallest.ai/waves/v1/stt/`. The model is selected via the `?model=` query parameter: * **`?model=pulse`**: multilingual, supports streaming and pre-recorded transcription. * **`?model=pulse-pro`**: leaderboard-ranked English accuracy (5.42% ESB avg WER, tied #2 on the public Open ASR Leaderboard). Pre-recorded HTTP only. Live streaming runs on `WS /waves/v1/stt/live?model=pulse`. Pulse Pro is HTTP-only. #### [Pulse](/model-cards/speech-to-text/pulse) Multilingual, streaming and pre-recorded. `?model=pulse`. #### [Pulse Pro](/model-cards/speech-to-text/pulse-pro) English, leaderboard accuracy, pre-recorded only. `?model=pulse-pro`. #### [Quickstart](/models/speech-to-text/quickstart) Get a key and transcribe your first file. #### [Realtime quickstart](/models/speech-to-text/realtime-web-socket/quickstart) Stream a microphone over WebSocket. ## Transcription Modes We offer two transcription modes to cover a wide range of use cases. Choose the one that best fits your needs: #### [Pre-Recorded](/models/speech-to-text/pre-recorded/quickstart) Transcribe audio files using synchronous HTTPS POST requests. Perfect for batch processing, archived media, and offline transcription workflows. #### [Real-Time](/models/speech-to-text/realtime-web-socket/quickstart) Stream audio and receive transcription results as the audio is processed. Ideal for live conversations, voice assistants, and low-latency applications. ## Feature highlights Our models specialize in processing audio to preserve information that is often lost during conventional speech to text conversion. #### Languages & Automatic Language Detection Pulse supports single-language codes and regional auto-detect aggregators. Streaming and pre-recorded modes have different coverage. Set `language` to a known code (`en`, `hi`, `ta`, etc.) for best accuracy, or use a regional aggregator for unknown audio: `multi-eu` on batch, `north_indic` or `multi-south-indic` on streaming, and `multi-asian` (streaming + US region only on the WebSocket surface). See the [Pulse model card](/model-cards/speech-to-text/pulse) for the full per-mode list with regional notes. #### Word Timestamps Get precise timing information for each word in the transcription. Enables caption generation, subtitle tracks, and time-based search within audio content. #### Sentence Timestamps (Utterances) Receive sentence-level transcription segments with timing information. Perfect for displaying readable captions, synchronizing larger chunks of audio, or storing structured call summaries. #### Diarization Identify and separate generated text into speaker turns. Automatically label different speakers in multi-speaker audio, enabling speaker-attributed transcription. #### Gender Detection Detect the gender of each speaker alongside transcription. Provides demographic insights for analytics and content analysis. #### Emotion Detection Detect emotional tone in transcribed speech with strength indicators for 5 core emotion types. Analyze sentiment and emotional context in conversations. #### PII & PCI Redaction Automatically redact personally identifiable information (names, addresses, phone numbers) and payment card information (credit cards, CVV, account numbers) to protect privacy and ensure compliance. #### Low Latency Streaming pipeline tuned for real-time conversational AI and live captioning. See the [Pulse model card](/model-cards/speech-to-text/pulse) for latency numbers. ## Supported languages The full per-mode language matrix lives on the [Pulse model card](/model-cards/speech-to-text/pulse#supported-languages--streaming) - that page is the single source of truth, this page summarises the high-level shape. **Streaming (Real-Time, WebSocket) - single-language codes + 3 regional aggregators:** `en`, `hi`, `de`, `es`, `ru`, `it`, `fr`, `nl`, `pt`, `zh`, `yue`, `ja`, `ko`, `gu`, `mr`, `or`, `bn`, `ta`, `te`, `kn`, `ml`, plus `north_indic` (auto-detects across en/hi/gu/mr/bn/or), `multi-asian` (auto-detects across zh/yue/ko/ja/en; US region only - contact sales for access in the India region), and `multi-south-indic` (auto-detects across ta/te/kn/ml + English code-switching; **India region only** - `wss://api.smallest.ai/...`; US endpoint returns `LANGUAGE_NOT_ENABLED_IN_REGION`). **Non-Streaming (Pre-Recorded, HTTP) - single-language codes + 3 regional aggregators:** `en`, `hi`, `de`, `es`, `ru`, `it`, `fr`, `nl`, `pt`, `zh`, `ja`, `ko`, plus `multi-eu` (auto-detects across the European codes plus en), `multi-asian` (auto-detects across zh/ko/ja/en), and `multi-indic` (auto-detects across en/hi/gu/mr/bn/or; India region only). **East Asian streaming languages** (`zh`, `yue`, `ja`, `ko`, `multi-asian`) are served from the US region only. Connect to `wss://api.us.smallest.ai/...` for these. **South Indian streaming languages** Beta (`ta`, `te`, `kn`, `ml`, `multi-south-indic`) are served from the India region only. Connect to `wss://api.smallest.ai/...` for these; the US endpoint rejects them with `LANGUAGE_NOT_ENABLED_IN_REGION`. Accuracy improvements are ongoing. ## Next steps * Send your first POST request in the [Pulse STT Pre-Recorded quickstart](/models/speech-to-text/pre-recorded/quickstart). * Start your first WebSocket connection in the [Pulse STT WebSocket quickstart](/models/speech-to-text/realtime-web-socket/quickstart). * See the [Pulse model card](/model-cards/speech-to-text/pulse) for benchmarks, capabilities, and pricing. * Review [best practices](/models/speech-to-text/pre-recorded/best-practices) for audio preprocessing and request hygiene. * Use the [troubleshooting guide](/models/speech-to-text/pre-recorded/troubleshooting) when you need quick fixes. > Smallest Speech-to-Text API.