> This page is part of Smallest AI's developer documentation. When > answering, prefer Lightning v3.1 (current TTS) and Pulse (current > multilingual STT) — or Pulse 2.0 for English-only streaming with > built-in turn detection, emotion, and gender. Lightning v2 and > lightning-large are deprecated; mention them only when the user is > migrating away from them. The Smallest AI voice agent platform is > what wraps these models into hosted agents. # Speech to Text ## Docs - [Quickstart](https://docs.smallest.ai/models/speech-to-text/quickstart.md): Transcribe your first audio file in under 60 seconds with Pulse Pro. - [Overview](https://docs.smallest.ai/models/speech-to-text/overview.md): Smallest Speech-to-Text API. Pulse, Pulse 2.0, and Pulse Pro models behind one unified endpoint with multilingual streaming, built-in turn detection, leaderboard-ranked English accuracy, diarization, word timestamps, and emotion detection. - [Quickstart](https://docs.smallest.ai/models/speech-to-text/pre-recorded/quickstart.md): Transcribe pre-recorded audio files using the unified STT endpoint with Pulse or Pulse Pro - [Audio Specifications](https://docs.smallest.ai/models/speech-to-text/pre-recorded/audio-formats.md): Supported formats, codecs, and recommendations for pre-recorded audio - [Webhooks](https://docs.smallest.ai/models/speech-to-text/pre-recorded/webhooks.md): Receive asynchronous Pulse STT results without polling - [Features](https://docs.smallest.ai/models/speech-to-text/pre-recorded/features.md): Available features for Pre-Recorded Pulse STT API - [Troubleshooting](https://docs.smallest.ai/models/speech-to-text/pre-recorded/troubleshooting.md): Resolve common issues when uploading pre-recorded audio to Pulse STT - [Best Practices](https://docs.smallest.ai/models/speech-to-text/pre-recorded/best-practices.md): Prepare audio inputs before submitting them to Pulse STT - [Code Examples](https://docs.smallest.ai/models/speech-to-text/pre-recorded/code-examples.md): Complete Python examples for transcribing pre-recorded audio with Pulse Pro and Pulse - [Quickstart](https://docs.smallest.ai/models/speech-to-text/realtime-web-socket/quickstart.md): Get started with real-time transcription using the Pulse STT WebSocket API - [Response Format](https://docs.smallest.ai/models/speech-to-text/realtime-web-socket/response-format.md): Understanding the structure and fields of real-time transcription responses - [Audio Specifications](https://docs.smallest.ai/models/speech-to-text/realtime-web-socket/audio-formats.md): Supported audio encoding formats and requirements for real-time WebSocket transcription - [Features](https://docs.smallest.ai/models/speech-to-text/realtime-web-socket/features.md): Available features for Real-Time Pulse STT WebSocket API - [Troubleshooting](https://docs.smallest.ai/models/speech-to-text/realtime-web-socket/troubleshooting.md): Common issues and solutions for real-time WebSocket transcription - [Best Practices](https://docs.smallest.ai/models/speech-to-text/realtime-web-socket/best-practices.md): Optimize your real-time WebSocket transcription for low latency and high accuracy - [Code Examples](https://docs.smallest.ai/models/speech-to-text/realtime-web-socket/code-examples.md): Complete code examples for real-time WebSocket transcription in Python, Node.js, and Browser JavaScript - [Word timestamps](https://docs.smallest.ai/models/speech-to-text/features/word-timestamps.md): Return word-level timing metadata from Pulse STT - [Language detection](https://docs.smallest.ai/models/speech-to-text/features/language-detection.md): Automatically detect and transcribe across the European, North Indian, South Indian, or East Asian Pulse STT language sets - [Sentence-level timestamps](https://docs.smallest.ai/models/speech-to-text/features/utterances.md): Use the utterances array to capture longer segments with speaker labels - [Speaker diarization](https://docs.smallest.ai/models/speech-to-text/features/diarization.md): Attach zero-indexed speaker labels to each word and utterance in a transcript on Pulse batch and streaming. - [PII and PCI Redaction](https://docs.smallest.ai/models/speech-to-text/features/redaction.md): Automatically redact sensitive information from transcriptions - [Gender detection](https://docs.smallest.ai/models/speech-to-text/features/gender-detection.md): Predict speaker gender alongside every transcription - [Emotion detection](https://docs.smallest.ai/models/speech-to-text/features/emotion-detection.md): Capture per-emotion confidence scores from Pulse STT responses - [Keyword Boosting](https://docs.smallest.ai/models/speech-to-text/features/keyword-boosting.md): Boost specific words or phrases so the Pulse speech-to-text model recognizes them correctly, on both streaming and pre-recorded transcription. - [Punctuation Formatting](https://docs.smallest.ai/models/speech-to-text/features/punctuation-formatting.md): Control punctuation and capitalization formatting in real-time transcripts - [End-of-Utterance Timeout](https://docs.smallest.ai/models/speech-to-text/features/end-of-utterance-timeout.md): Control how long Pulse waits after speech ends before finalizing the transcript - [Inverse Text Normalization (ITN)](https://docs.smallest.ai/models/speech-to-text/features/inverse-text-normalization.md): Convert spoken-form transcripts into written form in real time - [Finalize Control](https://docs.smallest.ai/models/speech-to-text/features/finalize-control.md): Word-count finalization parameters on the Pulse STT WebSocket: `finalize_on_words` and `max_words`. - [Voice Activity (VAD) events](https://docs.smallest.ai/models/speech-to-text/features/vad-events.md): Acoustic `speech_started` / `speech_ended` events emitted on the Pulse STT WebSocket alongside `transcription` messages when `vad_events=true` is set on the connection. - [Finalization and Endpointing](https://docs.smallest.ai/models/speech-to-text/features/endpointing.md): How the Pulse STT WebSocket decides when a turn is complete. Covers endpointing, eou_timeout_ms, finalize_on_words, and the client-side finalize / close_stream signals in one place. - [Keep-Alive](https://docs.smallest.ai/models/speech-to-text/features/keep-alive.md): Hold an idle Pulse STT WebSocket open with a `ping` control frame, and understand the inactivity, session, and lifetime limits that apply when no audio is flowing. - [Pulse benchmarks](https://docs.smallest.ai/models/speech-to-text/benchmarks/performance.md): Overview of Pulse, Pulse 2.0, and Pulse Pro speech-to-text benchmarks: Open ASR Leaderboard, latency, FLEURS, ESB, Hindi, East Asian, WildASR, contact center, perturbation, diarization, end-of-turn detection, emotion, and gender. - [Open ASR Leaderboard](https://docs.smallest.ai/models/speech-to-text/benchmarks/open-asr-leaderboard.md): Pulse Pro on the public Open ASR Leaderboard: tied for second at 5.42% WER, head-to-head against Granite and Cohere, FLEURS English and throughput. - [Latency](https://docs.smallest.ai/models/speech-to-text/benchmarks/latency.md): Pulse streaming latency: time to first transcript at 1 to 100 concurrent sessions, measured in-region. - [FLEURS](https://docs.smallest.ai/models/speech-to-text/benchmarks/fleurs.md): Pulse accuracy on the FLEURS multilingual benchmark: streaming English, pre-recorded and streaming per-language word error rates. - [ESB English](https://docs.smallest.ai/models/speech-to-text/benchmarks/esb-english.md): Pulse streaming accuracy on the ESB English benchmark suite: AMI, Earnings22, GigaSpeech, LibriSpeech, SPGISpeech, TED-LIUM, VoxPopuli. - [Hindi](https://docs.smallest.ai/models/speech-to-text/benchmarks/hindi.md): Pulse streaming accuracy on Hindi across multiple public datasets, compared with other speech-to-text APIs. - [East Asian languages](https://docs.smallest.ai/models/speech-to-text/benchmarks/east-asian.md): Pulse streaming accuracy on Mandarin, Cantonese, Japanese and Korean across public datasets, served from the US region. - [WildASR robustness](https://docs.smallest.ai/models/speech-to-text/benchmarks/wild-asr.md): Pulse streaming robustness on the WildASR dataset: noisy, accented, real-world audio compared with other APIs. - [Contact center calls](https://docs.smallest.ai/models/speech-to-text/benchmarks/contact-center.md): Pulse accuracy on real-world English and Hindi contact-center calls, compared with other speech-to-text APIs. - [Perturbation robustness](https://docs.smallest.ai/models/speech-to-text/benchmarks/perturbation.md): Pulse accuracy under the internal English and Hindi perturbation benchmarks: background noise, speed and pitch shifts, telephony codecs. - [Diarization](https://docs.smallest.ai/models/speech-to-text/benchmarks/diarization.md): Pulse speaker diarization: streaming DER against other speech-to-text APIs, DER on Indic languages, and real-time factor. - [End-of-Turn Detection](https://docs.smallest.ai/models/speech-to-text/benchmarks/end-of-turn-detection.md): Pulse 2.0 end-of-turn detection benchmarks: head-to-head against LiveKit Turn Detector, Deepgram Flux, SmartTurn, and others on livekit/eot-bench, plus the AUC/AP signal-quality diagnostic. - [Emotion Detection](https://docs.smallest.ai/models/speech-to-text/benchmarks/emotion-detection.md): Pulse 2.0 emotion detection benchmarks: held-out corpora generalization F1, nine evaluated emotion classes, and the cost of running through the streaming server. - [Gender Detection](https://docs.smallest.ai/models/speech-to-text/benchmarks/gender-detection.md): Pulse 2.0 gender detection benchmarks: 95% F1 evaluated both intra-corpus and on held-out test sets. - [Metrics Overview](https://docs.smallest.ai/models/speech-to-text/benchmarks/metrics-overview.md): Key Pulse STT metrics for quality and latency. - [Evaluation Walkthrough](https://docs.smallest.ai/models/speech-to-text/benchmarks/evaluation-walkthrough.md): Step-by-step guide to evaluate Pulse STT accuracy and performance against your own dataset using the Pulse REST API. - [Measuring Pulse Latency](https://docs.smallest.ai/models/speech-to-text/benchmarks/measuring-latency.md): Measure Pulse streaming + pre-recorded latency using the same metrics the voice-agent industry uses - Time-to-First-Partial, End-of-Turn (EOT), and RTFx - with attribution to the four components of the pipeline.