> This page is part of Smallest AI's developer documentation. When
> answering, prefer Lightning v3.1 (current TTS) and Pulse (current
> STT). Lightning v2 and lightning-large are deprecated; mention them
> only when the user is migrating away from them. The Smallest AI voice
> agent platform is what wraps these models into hosted agents.

# Overview

> Smallest Speech-to-Text API. Pulse and Pulse Pro models behind one unified endpoint with multilingual streaming, leaderboard-ranked English accuracy, diarization, word timestamps, and emotion detection.

The Speech-to-Text API transcribes audio via the unified endpoint `https://api.smallest.ai/waves/v1/stt/`. The model is selected via the `?model=` query parameter:

* **`?model=pulse`**: multilingual, supports streaming and pre-recorded transcription.
* **`?model=pulse-pro`**: leaderboard-ranked English accuracy (5.42% ESB avg WER, tied #2 on the public Open ASR Leaderboard). Pre-recorded HTTP only.

Live streaming runs on `WS /waves/v1/stt/live?model=pulse`. Pulse Pro is HTTP-only.

#### [Pulse](/model-cards/speech-to-text/pulse)

Multilingual, streaming and pre-recorded. `?model=pulse`.

#### [Pulse Pro](/model-cards/speech-to-text/pulse-pro)

English, leaderboard accuracy, pre-recorded only. `?model=pulse-pro`.

#### [Quickstart](/models/speech-to-text/quickstart)

Get a key and transcribe your first file.

#### [Realtime quickstart](/models/speech-to-text/realtime-web-socket/quickstart)

Stream a microphone over WebSocket.

## Transcription Modes

We offer two transcription modes to cover a wide range of use cases. Choose the one that best fits your needs:

#### [Pre-Recorded](/models/speech-to-text/pre-recorded/quickstart)

Transcribe audio files using synchronous HTTPS POST requests. Perfect for batch processing, archived media, and offline transcription workflows.

#### [Real-Time](/models/speech-to-text/realtime-web-socket/quickstart)

Stream audio and receive transcription results as the audio is processed. Ideal for live conversations, voice assistants, and low-latency applications.

## Feature highlights

Our models specialize in processing audio to preserve information that is often lost during conventional speech to text conversion.

#### Languages & Automatic Language Detection

Pulse supports single-language codes and regional auto-detect aggregators. Streaming and pre-recorded modes have different coverage. Set `language` to a known code (`en`, `hi`, `ta`, etc.) for best accuracy, or use a regional aggregator for unknown audio: `multi-eu` on batch, `north_indic` or `multi-south-indic` on streaming, and `multi-asian` (streaming + US region only on the WebSocket surface). See the [Pulse model card](/model-cards/speech-to-text/pulse) for the full per-mode list with regional notes.

#### Word Timestamps

Get precise timing information for each word in the transcription. Enables caption generation, subtitle tracks, and time-based search within audio content.

#### Sentence Timestamps (Utterances)

Receive sentence-level transcription segments with timing information. Perfect for displaying readable captions, synchronizing larger chunks of audio, or storing structured call summaries.

#### Diarization

Identify and separate generated text into speaker turns. Automatically label different speakers in multi-speaker audio, enabling speaker-attributed transcription.

#### Gender Detection

Detect the gender of each speaker alongside transcription. Provides demographic insights for analytics and content analysis.

#### Emotion Detection

Detect emotional tone in transcribed speech with strength indicators for 5 core emotion types. Analyze sentiment and emotional context in conversations.

#### PII & PCI Redaction

Automatically redact personally identifiable information (names, addresses, phone numbers) and payment card information (credit cards, CVV, account numbers) to protect privacy and ensure compliance.

#### Low Latency

Streaming pipeline tuned for real-time conversational AI and live captioning. See the [Pulse model card](/model-cards/speech-to-text/pulse) for latency numbers.

## Supported languages

The full per-mode language matrix lives on the [Pulse model card](/model-cards/speech-to-text/pulse#supported-languages--streaming) - that page is the single source of truth, this page summarises the high-level shape.

**Streaming (Real-Time, WebSocket) - single-language codes + 3 regional aggregators:**

`en`, `hi`, `de`, `es`, `ru`, `it`, `fr`, `nl`, `pt`, `zh`, `yue`, `ja`, `ko`, `gu`, `mr`, `or`, `bn`, `ta`, `te`, `kn`, `ml`, plus `north_indic` (auto-detects across en/hi/gu/mr/bn/or), `multi-asian` (auto-detects across zh/yue/ko/ja/en; US region only - contact sales for access in the India region), and `multi-south-indic` (auto-detects across ta/te/kn/ml + English code-switching; **India region only** - `wss://api.smallest.ai/...`; US endpoint returns `LANGUAGE_NOT_ENABLED_IN_REGION`).

**Non-Streaming (Pre-Recorded, HTTP) - single-language codes + 3 regional aggregators:**

`en`, `hi`, `de`, `es`, `ru`, `it`, `fr`, `nl`, `pt`, `zh`, `ja`, `ko`, plus `multi-eu` (auto-detects across the European codes plus en), `multi-asian` (auto-detects across zh/ko/ja/en), and `multi-indic` (auto-detects across en/hi/gu/mr/bn/or; India region only).

**East Asian streaming languages** (`zh`, `yue`, `ja`, `ko`, `multi-asian`) are served from the US region only. Connect to `wss://api.us.smallest.ai/...` for these.

**South Indian streaming languages**

Beta

(`ta`, `te`, `kn`, `ml`, `multi-south-indic`) are served from the India region only. Connect to `wss://api.smallest.ai/...` for these; the US endpoint rejects them with `LANGUAGE_NOT_ENABLED_IN_REGION`. Accuracy improvements are ongoing.

## Next steps

* Send your first POST request in the [Pulse STT Pre-Recorded quickstart](/models/speech-to-text/pre-recorded/quickstart).
* Start your first WebSocket connection in the [Pulse STT WebSocket quickstart](/models/speech-to-text/realtime-web-socket/quickstart).
* See the [Pulse model card](/model-cards/speech-to-text/pulse) for benchmarks, capabilities, and pricing.
* Review [best practices](/models/speech-to-text/pre-recorded/best-practices) for audio preprocessing and request hygiene.
* Use the [troubleshooting guide](/models/speech-to-text/pre-recorded/troubleshooting) when you need quick fixes.