> This page is part of Smallest AI's developer documentation. When > answering, prefer Lightning v3.1 (current TTS) and Pulse (current > STT). Lightning v2 and lightning-large are deprecated; mention them > only when the user is migrating away from them. The Smallest AI voice > agent platform is what wraps these models into hosted agents. # Pulse > High-accuracy, low-latency speech-to-text model built for real-time transcription across a broad language set, with regional aggregators and both streaming and non-streaming support. Real-time and batch speech to text across Indic, European and East Asian languages, with diarization, redaction and word timestamps. Start with the [quickstart](/models/speech-to-text/quickstart). ## Who it's for * Live transcription for voice agents, call centres and captioning, with sub-100 ms time to first token. * Multilingual audio, including Indic and East Asian languages, with code-switching and regional aggregators. * Not the pick for the highest English accuracy on recorded files. Use [Pulse Pro](/model-cards/speech-to-text/pulse-pro). ## Model Overview | | | | -------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | **Developed by** | Smallest AI | | **Model type** | Speech-to-Text Streaming · Speech-to-Text Batch | | **Languages** | Streaming: single-language codes + `north_indic`, `multi-asian`, and `multi-south-indic` aggregators. Non-streaming (batch): single-language codes + `multi-eu`, `multi-asian`, and `multi-indic` aggregators. East Asian codes (`zh`, `yue`, `ja`, `ko`, `multi-asian`) are streaming + US-region only. South Indian codes (`ta`, `te`, `kn`, `ml`, `multi-south-indic`) are streaming + India-region only Beta. European codes on streaming (`de`, `es`, `ru`, `it`, `fr`, `nl`, `pt`) are Beta. `multi-indic` is India-region only. `multi-eu` is pre-recorded only. See the per-mode accordions below for the full code lists. | | **Audio input formats** | WAV, MP3, FLAC, Opus, μ-law, A-law, raw PCM | | **Pricing** (Standard Plan) | Realtime: \$0.004/min · Batch: \$0.003/min | | **Concurrency** (Standard Plan) | Streaming: 100 concurrent requests · Batch: 25 RPM | | **Recommended Sample Rate** | 16,000 Hz | | **Recommended GPU** | 1× NVIDIA L4 (24 GB VRAM). Larger GPUs (L40S, A100, H100) supported. | --- ## Key Capabilities Sub-100 ms TTFT at 1 concurrency, \~300 ms at 100 concurrent requests. Designed for live transcription and conversational AI. Single-language codes with regional auto-detect aggregators and in-session code-switching. Streaming and pre-recorded modes have different coverage; see [Supported Languages](#supported-languages). Built-in redaction of personal data and payment-card information on both streaming and non-streaming surfaces. Automatic multi-speaker identification with per-word and per-utterance speaker labels. Background-noise handling built into the model - no preprocessing required. Multi-language audio within a single session. Set the known primary language (e.g. `es` for Spanish) - English+Spanish is handled automatically. --- ## Performance & Benchmarks #### [#1 for Speed Factor](https://artificialanalysis.ai/speech-to-text/non-streaming?speed=speedFactor) Artificial Analysis #### [3.2 WER](https://benchmarks.coval.ai/stt/) Coval Benchmarks #### [#1 for P95 Latency](https://research.sierra.ai/mubench/) MuBench by Sierra Pulse STT is evaluated against three open-source datasets - [FLEURS](https://huggingface.co/datasets/google/fleurs), [ESB](https://huggingface.co/datasets/esb/datasets), and [WildASR](https://huggingface.co/datasets/bosonai/WildASR) - and one internal English perturbation suite. Word Error Rate (WER) by language. Lower is better. `NA` = not available or not supported by that provider. For the full benchmark comparison across every dataset, see the [Performance page](/models/speech-to-text/benchmarks/performance). Pick the benchmark closest to your workload - each accordion below expands its full table. #### FLEURS Dataset Streaming - English WER on the English subset of FLEURS across providers in streaming mode. Lower is better. | Provider | Smallest Pulse | Assembly Universal 3 Pro | AWS transcribe | Azure | Deepgram Nova 3 | Grok | Sarvam Saras 3 | ElevenLabs Scribe V2 | | :------- | :------------: | :----------------------: | :------------: | :----: | :-------------: | :----: | :------------: | :------------------: | | **WER** | 5.55% | 3.13% | 6.54% | 13.79% | 11.59% | 60.00% | 6.34% | 3.88% | #### A note on audio amplitude normalization Audio amplitude normalization materially changes WER on FLEURS. Most competitors benchmark on raw FLEURS - which has variable, often low amplitude - without normalizing peak audio to −10 dBFS. This makes some models look much better than they actually are. Pulse is stable across all amplitude regimes. | Model | Raw FLEURS | −10 dBFS | −20 dBFS | Stable across regimes? | | :------------------ | :--------: | :------: | :------: | :-------------------------------- | | **Pulse** | 5.55% | 6.06% | 5.81% | Yes | | **Deepgram Nova 3** | 11.59% | 6.57% | 6.51% | Partial - 1.8× degradation on raw | | **Grok** | 60.00% | 7.58% | 8.59% | Collapses on raw | #### FLEURS Dataset Streaming - European + Indic Languages WER on FLEURS in streaming mode, broken down by language family. Lower is better. #### European Languages | Language | Smallest Pulse | Deepgram Nova 2 | Deepgram Nova 3 | | -------------- | -------------- | --------------- | --------------- | | **Italian** | **4.41%** | 11.05% | 6.99% | | **English** | **5.55%** | 15.59% | 11.21% | | **Spanish** | **5.99%** | 10.67% | 7.52% | | **Portuguese** | **8.32%** | 14.15% | 11.46% | | **German** | **9.5%** | 11.1% | 10.15% | | **French** | **10.71%** | 14.3% | 12.07% | | **Russian** | **14.35%** | NA | NA | | **Dutch** | **11.90%** | NA | NA | #### Indic Languages | Language | Smallest Pulse | Deepgram Nova 2 | Deepgram Nova 3 | | ------------ | :------------: | :-------------: | :-------------: | | **Hindi** | **8.3%** | 20.0% | 15.46% | | **Marathi** | **15.68%** | NA | NA | | **Gujarati** | **20.05%** | NA | NA | #### FLEURS Dataset Batch - European + Indic Languages WER on FLEURS in pre-recorded mode (full-file upload). Lower is better. #### European Languages | Language | Smallest Pulse | Deepgram Nova 2 | Deepgram Nova 3 | | -------------- | -------------- | --------------- | --------------- | | **English** | **4.55%** | 7.9% | 6.7% | | **Italian** | **3.0%** | 10.7% | 6.2% | | **Spanish** | **3.2%** | 8.6% | 4.1% | | **Portuguese** | **5.0%** | 9.9% | 7.5% | | **German** | **6.4%** | 8.2% | 8.5% | | **French** | **7.1%** | 13.3% | 10.7% | | **Russian** | 9.6% | 7.9% | 11.8% | | **Dutch** | 15.0% | 16.3% | 12.5% | #### Indic Languages | Language | Smallest Pulse | Deepgram Nova 2 | Deepgram Nova 3 | | --------- | :------------: | :-------------: | :-------------: | | **Hindi** | **6.3%** | 23.5% | 23.6% | #### VISTAAR Streaming - Hindi WER across seven Hindi datasets covering read speech, conversational speech, telephony / contact-center audio, and noise-augmented variants. Compared against IndicWhisper, Sarvam Saaras v3, scribe v2, and Deepgram Nova-3. Lower is better. | Dataset | Smallest Pulse | IndicWhisper | Sarvam Saaras v3 | scribe v2 | Deepgram Nova-3 | | -------------------- | :------------: | :----------: | :--------------: | :-------: | :-------------: | | **FLEURS** | **6.17** | 15.00 | 6.18 | 8.96 | 14.09 | | **Kathbath** | **5.68** | 10.30 | 6.10 | 8.67 | 16.22 | | **Kathbath (noisy)** | 6.95 | 12.00 | 5.10 | 10.11 | 17.06 | | **Common Voice** | **7.62** | 11.40 | 10.36 | 13.61 | 23.55 | | **Indic-TTS** | **3.63** | 7.60 | 5.16 | 8.75 | 10.72 | | **MUCS** | **3.22** | 12.00 | 4.69 | 8.15 | 16.20 | | **Gramvaani** | **15.62** | 26.80 | 20.80 | 24.09 | 31.44 | For the full breakdown including training-data and evaluation-protocol notes, see the [Performance page](/models/speech-to-text/benchmarks/performance#hindi-multi-dataset-streaming). #### HF ESB Dataset Streaming - English A Hugging Face benchmark suite aggregating 9 English speech datasets across diverse domains (audiobooks, parliament, meetings, finance, etc.) to test STT generalization. Lower WER is better. *Evaluated on the open-source Hugging Face ESB datasets. Numbers from internal evaluation.* | Dataset | Smallest Pulse | Assembly Universal 3 Pro | AWS Transcribe | Azure | Deepgram Nova 3 | Grok | Sarvam Saras V3 | ElevenLabs Scribe V2 | | :-------------------- | :------------: | :----------------------: | :------------: | :---: | :-------------: | :---: | :-------------: | :------------------: | | **LibriSpeech Clean** | 2.11 | 1.65 | 2.16 | 2.48 | 3.20 | 3.61 | 3.09 | 1.97 | | **LibriSpeech Other** | 4.54 | 2.86 | 4.88 | 5.74 | 6.60 | 7.28 | 6.85 | 4.45 | | **Common Voice** | 12.29 | 6.73 | 10.69 | 47.28 | 14.22 | 43.46 | 11.37 | 9.83 | | **VoxPopuli** | 7.06 | 7.28 | 7.07 | 14.10 | 9.55 | 11.49 | 7.77 | 7.91 | | **TED-LIUM** | 2.52 | 2.95 | 2.66 | 3.81 | 3.59 | 6.90 | 2.89 | 3.16 | | **GigaSpeech** | 9.68 | 9.12 | 10.09 | 5.35 | 10.05 | 10.05 | 9.57 | 9.66 | | **SPGISpeech** | 2.41 | 1.74 | 4.18 | 3.53 | 2.99 | 9.70 | 3.89 | 4.40 | | **Earnings22** | 12.25 | 11.52 | 12.21 | 8.54 | 15.79 | 27.02 | 11.97 | 12.20 | | **AMI** | 10.04 | 14.60 | 13.19 | 8.46 | 17.04 | 19.19 | 13.08 | 12.23 | | **Aggregate** | 6.99 | 6.49 | 7.46 | 11.03 | 9.23 | 15.41 | 7.83 | 7.31 | #### WildASR Dataset Streaming - English (STT Robustness) An open-source robustness benchmark designed to stress-test STT under real-world degraded conditions: clipping, far-field capture, background noise, phone codec compression, reverberation, and accented speech. Lower WER is better. `n/a` = not supported by that provider. *Evaluated on the open-source WildASR dataset. Numbers from internal evaluation.* | Dataset | Smallest Pulse | Assembly Universal 3 pro | AWS Transcribe | Azure | Deepgram Nova 3 | Sarvam Saras V3 | ElevenLabs Scribe | | :---------------- | :------------: | :----------------------: | :------------: | :---: | :-------------: | :-------------: | :---------------: | | **Clean** | 5.31 | 3.33 | 7.01 | 11.11 | 11.62 | 7.02 | 4.24 | | **Clipping** | 10.31 | 6.59 | 42.10 | 4.35 | 47.35 | 28.74 | 11.20 | | **Far-field** | 8.99 | 26.07 | 38.76 | n/a | 62.99 | 21.27 | 7.38 | | **Noise Gap** | 6.91 | 4.04 | 9.77 | n/a | 15.04 | 9.74 | 6.30 | | **Phone Codec** | 6.64 | 3.45 | 8.70 | n/a | 9.13 | 10.64 | 4.98 | | **Reverberation** | 8.06 | 23.50 | 14.83 | n/a | 27.27 | 4.35 | 6.48 | | **Accent** | 5.74 | 2.80 | 4.45 | n/a | 7.31 | n/a | 4.01 | | **Aggregate** | 7.60 | 12.52 | 18.35 | 8.82 | 28.17 | 17.75 | 6.47 | #### Real-World Contact Center Calls (English & Hindi) WER and Character Error Rate (CER) on live, dual-channel contact center call recordings (agent and customer channels scored separately), covering English and Hindi. Diarization was applied first to cleanly separate agent and customer speech before scoring. Lower is better. *Evaluated on real-world telephony call audio. Numbers from internal evaluation.* #### English | STT Model | WER % | CER % | | :----------------- | :-------: | :------: | | **Smallest Pulse** | **11.47** | **8.18** | | Sarvam Saaras v4 | 12.18 | 9.20 | | Deepgram Nova 3 | 15.80 | 10.99 | | Soniox v5 | 16.26 | 11.64 | #### Hindi | STT Model | WER % | CER % | | :----------------- | :-------: | :------: | | **Smallest Pulse** | **10.20** | **6.74** | | Sarvam Saaras v4 | 12.04 | 8.78 | | Deepgram Nova 3 | 13.07 | 10.35 | | Soniox v5 | 15.82 | 12.45 | Pulse leads across both languages on real production call audio, extending its edge on Hindi CER in particular. For the full breakdown, see the [Performance page](/models/speech-to-text/benchmarks/performance#real-world-contact-center-calls-english--hindi). #### Multi-dataset Streaming - East Asian Languages WER for the four East Asian languages newly enabled on the streaming endpoint (us-west-2). Three datasets per language covering read speech (FLEURS), conversational/crowdsourced speech (Common Voice 25), and language-specific corpora (JSUT, Zeroth-Korean, MDCC, AISHELL-1). Compared head-to-head against Deepgram Nova-3. Lower WER is better. | Lang | Dataset | Smallest Pulse | Deepgram Nova 3 | | :------------ | :------------- | :------------: | :-------------: | | **Japanese** | CV-25 | **23.84%** | 34.81% | | **Japanese** | FLEURS | **10.78%** | 17.11% | | **Japanese** | JSUT BASIC5000 | **11.47%** | 11.65% | | **Korean** | CV-25 | 9.79% | **9.66%** | | **Korean** | FLEURS | **7.95%** | 10.79% | | **Korean** | Zeroth-Korean | **5.25%** | 6.46% | | **Cantonese** | CV-25 | **6.16%** | 14.09% | | **Cantonese** | FLEURS | **13.06%** | 15.43% | | **Cantonese** | MDCC | **5.85%** | 12.77% | | **Mandarin** | CV-25 | **15.99%** | 22.44% | | **Mandarin** | FLEURS | 14.25% | **13.89%** | | **Mandarin** | AISHELL-1 | **7.34%** | 8.69% | | **Average** | - | **10.91%** | 15.50% | Pulse averages **10.91% WER** vs Deepgram Nova-3's **15.50%** across the four East Asian languages - Pulse leads on 10 of 12 dataset rows, with the largest gains on Japanese CV-25 and Cantonese CV-25. These four languages stream from `wss://api.us.smallest.ai/waves/v1/stt/live?model=pulse` only (US region). See the [streaming Asian-language documentation](/models/speech-to-text/overview) for the region-routing details. #### Internal Perturbation Benchmark Streaming - English Not a public dataset. The English audio is sliced by perturbation type (Noise, Silence, Telephony 911, Boundary, Disfluency, Long Audios, Repetition, Entity, Accent, Emotion, Speaker Diversity, Speed, Pitch, Volume, Audio Quality) to isolate model weaknesses. Lower WER is better. | Category | Smallest Pulse | Assembly Universal 3 Pro | AWS Transcribe | Deepgram Nova 3 | ElevenLabs Scribe | | :-------------------- | :------------: | :----------------------: | :------------: | :-------------: | :---------------: | | **Noise** | 10.96 | 11.93 | 14.19 | 14.58 | 10.05 | | **Silence** | 7.18 | 4.22 | 8.22 | 13.28 | 10.61 | | **Telephony 911** | 21.05 | 23.93 | 27.88 | 28.43 | 20.29 | | **Boundary** | 2.74 | 3.09 | 3.18 | 3.66 | 1.73 | | **Disfluency** | 8.20 | 7.81 | 9.23 | 8.62 | 9.29 | | **Long Audios** | 6.60 | 8.58 | 11.66 | 11.16 | 9.25 | | **Repetition** | 9.10 | 9.82 | 10.39 | 9.57 | 10.81 | | **Entity** | 6.69 | 10.13 | 13.35 | 11.69 | 9.48 | | **Accent** | 8.27 | 7.89 | 9.51 | 10.42 | 7.25 | | **Emotion** | 13.53 | 16.34 | 18.57 | 18.07 | 11.84 | | **Speaker Diversity** | 6.90 | 6.72 | 8.81 | 9.48 | 5.95 | | **Speed** | 3.67 | 3.63 | 4.40 | 6.88 | 3.74 | | **Pitch** | 2.89 | 3.07 | 3.21 | 4.07 | 1.61 | | **Volume** | 2.43 | 3.05 | 2.41 | 3.67 | 1.47 | | **Audio Quality** | 2.69 | 2.86 | 3.03 | 4.08 | 1.60 | | **Average WER** | 7.53 | 8.20 | 9.87 | 10.51 | 7.66 | #### Internal Perturbation Benchmark Streaming - Hindi Not a public dataset. Hindi audio sliced by perturbation type to isolate model weaknesses. Lower WER is better except for Entity EDR where higher is better (↑). | Category | Smallest Pulse | Sarvam Saras V3 | Deepgram Nova 3 | | :----------------- | :------------: | :-------------: | :-------------: | | **Noise** | **15.76%** | 22.18% | 21.52% | | **Silence** | **8.22%** | 11.38% | 18.40% | | **Entity** | **8.27%** | 17.36% | 14.67% | | **Entity NE-WER** | **13.32%** | 26.72% | 26.58% | | **Entity EDR (↑)** | **83.13%** | 76.13% | 67.80% | | **Boundary** | **8.03%** | 17.52% | 17.36% | | **Long Audios** | **12.00%** | 18.42% | 19.21% | | **Speed** | **14.77%** | 21.39% | 38.21% | | **Pitch** | **8.14%** | 11.92% | 19.59% | | **Audio Quality** | **10.86%** | 11.75% | 19.51% | | **Volume** | **7.08%** | 15.25% | 16.76% | | **Disfluency** | **10.77%** | 12.06% | 18.44% | | **Repetition** | **8.11%** | 11.27% | 20.40% | #### Diarization - accuracy and speed Pulse ships speaker diarization on both batch and streaming through the same API — no separate diarization vendor to stitch in. See the [Diarization](/models/speech-to-text/features/diarization) page for the response shape and enable flag. **Metric.** DER (Diarization Error Rate) = percent of speech time handled wrong (wrong speaker + missed speech + falsely detected speech). Scored with no collar, overlapping speech counted, permutation-optimal speaker matching, duration-pooled via `pyannote.metrics`. Lower is better. #### Streaming DER vs other STT APIs Per-dataset DER. Lower is better. `-` means we have not scored ElevenLabs on that dataset yet. | Dataset | Pulse | ElevenLabs | Deepgram | | :----------------------------- | :-------: | :--------: | :------: | | ami\_ihm (headset meeting) | **15.28** | 33.6 | 29.03 | | ami\_sdm (far-field meeting) | **23.13** | 37.6 | 37.86 | | dihard3\_farfield (clean) | **33.18** | - | 47.69 | | sbcsae (clean) | **31.97** | - | 36.77 | | callhome\_eng (English phone) | **14.94** | 23.0 | 27.71 | | callhome\_deu (German phone) | **14.9** | 23.9 | 25.8 | | callhome\_spa (Spanish phone) | **31.0** | 39.6 | 41.4 | | simsamu (French emergency 911) | **14.2** | 16.3 | 38.1 | | **macro** | **23.70** | - | 35.81 | Extended benchmark against the specialist API pyannoteAI (diarization-only, not an STT API): pyannoteAI wins on some meeting-room datasets (`ami_ihm` 12.8 vs 18.5, `voxconverse` 7.7 vs 20.8) where speaker-count is high. Pulse remains competitive and wins on French emergency dispatch (`simsamu` 13.4 vs 14.7). #### Streaming DER on Indic languages Streaming diarization measured end-to-end through the WebSocket API at concurrency 1. Lower is better. Pulse macro-average across all 22 Indic language codes is 23.2 DER; Pyannote Live-1 is 37.1; Deepgram Live is 77.6 (Deepgram Live does not have a first-class Indic pipeline). #### Speed (real-time factor) RTF = processing time / audio length. Lower is faster. `1/RTF` is "times faster than real-time." Measured **through each vendor's API at concurrency 1 with diarization enabled** — the round-trip a customer actually sees, not local compute. | Vendor | RTF (median) | Times real-time | | :--------- | :----------: | :-------------: | | **Pulse** | **\~0.004** | **\~240×** | | Deepgram | \~0.007 | \~140× | | pyannoteAI | \~0.022 | \~45× | | AssemblyAI | \~0.028 | \~36× | | ElevenLabs | \~0.059 | \~17× | Streaming diarization emits speaker labels on `is_final: true` transcription events with roughly the same latency as transcription itself (\~1 s tail from utterance end to final event on typical connectivity). --- ## Supported Languages Pulse supports single-language codes plus regional auto-detect aggregators. Coverage differs between streaming and pre-recorded modes. Click an accordion to expand the full per-mode list. #### Streaming (single-language codes + 3 regional aggregators) | # | Language | Code | | -- | ------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | 1 | English | `en` | | 2 | Hindi (bilingual, mixed-script) | `hi` | | 3 | Hindi (Devanagari-only) | `hi-dev` (India region only \[\*\*]) | | 4 | German Beta | `de` | | 5 | Spanish Beta | `es` | | 6 | Russian Beta | `ru` | | 7 | Italian Beta | `it` | | 8 | French Beta | `fr` | | 9 | Dutch Beta | `nl` | | 10 | Portuguese Beta | `pt` | | 11 | Mandarin | `zh` (US region only \[\*]) | | 12 | Cantonese | `yue` (US region only \[\*]) | | 13 | Japanese | `ja` (US region only \[\*]) | | 14 | Korean | `ko` (US region only \[\*]) | | 15 | Gujarati | `gu` | | 16 | Marathi | `mr` | | 17 | Oriya | `or` | | 18 | Bengali | `bn` | | 19 | Tamil Beta | `ta` (India region only \[\*\*]) | | 20 | Telugu Beta | `te` (India region only \[\*\*]) | | 21 | Kannada Beta | `kn` (India region only \[\*\*]) | | 22 | Malayalam Beta | `ml` (India region only \[\*\*]) | | 23 | North-Indic aggregator | `north_indic` (auto-detects across `en`, `hi`, `gu`, `mr`, `bn`, `or`) | | 24 | Multi-Asian aggregator | `multi-asian` (auto-detects across `zh`, `yue`, `ko`, `ja`, `en`). US region only \[\*]. Contact sales for access in the India region. | | 25 | South-Indic aggregator Beta | `multi-south-indic` (auto-detects across `ta`, `te`, `kn`, `ml`, and English code-switching). India region only \[\*\*]. Use when the South Indian language is not known in advance. | \[\*] East Asian languages (`zh`, `yue`, `ja`, `ko`, `multi-asian`) are served from the US region only. Connect to `wss://api.us.smallest.ai/waves/v1/stt/live` instead of `wss://api.smallest.ai/...` for these. \[\*\*] South Indian languages (`ta`, `te`, `kn`, `ml`, `multi-south-indic`) and `hi-dev` are served from the India region only (`wss://api.smallest.ai/...`). Requests to `wss://api.us.smallest.ai/...` are rejected with error code `LANGUAGE_NOT_ENABLED_IN_REGION`. Contact support to request access. #### Non-Streaming (Pre-Recorded, single-language codes + 3 regional aggregators) | # | Language | Code | | -- | ---------------------- | ------------------------------------------------------------------------------------------ | | 1 | English | `en` | | 2 | Hindi | `hi` | | 3 | German | `de` | | 4 | Spanish | `es` | | 5 | Russian | `ru` | | 6 | Italian | `it` | | 7 | French | `fr` | | 8 | Dutch | `nl` | | 9 | Portuguese | `pt` | | 10 | Mandarin | `zh` | | 11 | Japanese | `ja` | | 12 | Korean | `ko` | | 13 | Multi-EU aggregator | `multi-eu` (auto-detects across the European codes above plus `en`). Pre-recorded only. | | 14 | Multi-Asian aggregator | `multi-asian` (auto-detects across `zh`, `ko`, `ja`, `en`) | | 15 | Multi-Indic aggregator | `multi-indic` (auto-detects across `en`, `hi`, `gu`, `mr`, `bn`, `or`). India region only. | > **Note** > > **Single language code vs. aggregator:** Use a specific language code (e.g. `hi`, `es`, `en`) whenever you know the language of the audio - the model optimizes directly for that language and also handles code-switching with English (e.g. `hi` covers Hindi–English mixed speech). Use an aggregator (`north_indic`, `multi-eu`, `multi-asian`) only when the language is genuinely unknown or the source is mixed across multiple languages; auto-detection adds a small accuracy overhead compared to an explicit code. --- ## Features - Streaming | Feature | Notes | | ------------------------- | ----------------------------------------- | | Speaker diarization | Identifies and labels each speaker | | Keyword boosting | Improves accuracy for custom vocabulary | | PII redaction | Personal information redaction | | PCI redaction | Payment card data redaction | | Word-level timestamps | Start and end time for each word | | Sentence-level timestamps | Start and end time for each sentence | | Punctuation | Automatically adds punctuation | | Code-switching | Handles multiple languages in one session | ## Features - Non-streaming | Feature | Notes | | ------------------------- | ----------------------------------------- | | Speaker diarization | Identifies and labels each speaker | | Keyword boosting | Improves accuracy for custom vocabulary | | PII redaction | Personal information redaction | | PCI redaction | Payment card data redaction | | Word-level timestamps | Start and end time for each word | | Sentence-level timestamps | Requires `word_timestamps=true` | | Punctuation | Automatically adds punctuation | | Code-switching | Handles multiple languages in one session | --- ## API Reference | Endpoint | Method | Use case | | ----------------------------------------------------- | --------- | -------------------------- | | `https://api.smallest.ai/waves/v1/stt/?model=pulse` | POST | Pre-recorded transcription | | `wss://api.smallest.ai/waves/v1/stt/live?model=pulse` | WebSocket | Streaming transcription | See [Transcribe (Pre-recorded)](/api-reference/models/speech-to-text/transcribe) for the full request/response schema, supported parameters, and error codes. The streaming surface shares parameters where applicable; see the [Realtime quickstart](/models/speech-to-text/realtime-web-socket/quickstart) for the WebSocket protocol details. --- ## Throughput, Latency & Pricing | Mode | Typical | Notes | | -------------------------------- | ------------- | ------------------------------------------------------------------------------------------------------ | | Streaming TTFT (1 concurrency) | 150 ms | See [Measuring latency](/models/speech-to-text/benchmarks/measuring-latency) for methodology. | | Streaming TTFT (100 concurrency) | \~300 ms | Concurrency curve documented in the [Performance page](/models/speech-to-text/benchmarks/performance). | | Pre-recorded RTFx | 50× or higher | Wall-clock; `metadata.processing_time_ms` excludes network. | Rate limits, concurrency caps, and pricing tiers are documented on the [Concurrency & Limits](/api-reference/concurrency-and-limits) page. For enterprise pricing, contact [sales@smallest.ai](mailto:sales@smallest.ai). --- ## Use Cases | Direct Use | Downstream Use | | ----------------------------------- | -------------------------------- | | Real-time call transcription | Multi-turn conversational agents | | Voice assistant input | Voice-to-text pipelines | | Meeting transcription | Telephony and IVR systems | | Accessibility and captioning | Content indexing and search | | Customer support recording analysis | Compliance and audit logging | --- ## FAQ #### What is the difference between Pulse streaming and batch modes? Pulse runs two independent inference paths. Streaming uses a WebSocket connection and emits partial transcripts in real time (suited for live call transcription, voice assistants, and conversational AI where first-token latency matters). Batch accepts a full audio file over HTTP and returns the complete transcript once processing is done (suited for call recordings, media archives, and any workload where you have the full audio upfront). Both modes support keyword boosting, diarization, word-level timestamps, and PII/PCI redaction. Streaming covers 21 languages; batch covers 12 languages, all of which are also on streaming. #### How do I choose between a specific language code and a regional aggregator? Use a specific language code (e.g. `hi`, `es`, `en`) whenever you know the language of the audio. Pulse optimises directly for that language and handles code-switching with English automatically, so `hi` covers Hindi/English mixed speech without needing an aggregator. Use `north_indic` or `multi-asian` only when the source language is genuinely unknown or mixed across several languages. Auto-detection adds a small accuracy overhead compared to an explicit code. #### Where can I find the full API reference and quickstart guides? The complete request/response schema, all query parameters, error codes, and WebSocket protocol details are in the [API Reference](/api-reference/models/speech-to-text/transcribe). For step-by-step setup, see the [Realtime quickstart](/models/speech-to-text/realtime-web-socket/quickstart) for streaming and the [Pre-recorded quickstart](/models/speech-to-text/pre-recorded/quickstart) for batch. #### What is the difference between Pulse and Pulse Pro? Pulse supports both streaming and batch transcription across a broad multilingual set with regional aggregators. Pulse Pro is English-only and batch-only, but achieves higher accuracy: tied for #2 on the Open ASR Leaderboard at 5.42% average WER vs Pulse's 5.55% on English FLEURS. Use Pulse for live streaming, multilingual audio, or latency-sensitive workloads. Use [Pulse Pro](/model-cards/speech-to-text/pulse-pro) for high-volume English batch transcription where accuracy is the top requirement. #### How does PII and PCI redaction work? Pulse applies built-in redaction on both streaming and batch surfaces (no preprocessing required). PII redaction masks personal identifiers such as names, phone numbers, email addresses, and SSNs. PCI redaction masks payment card data including card numbers, CVVs, and expiry dates. Both are enabled via query parameters in the API request. Redacted content is replaced with a placeholder in the transcript. The original audio is not retained post-processing. Supported on `language=en` and `language=hi` only. --- ## Safety & Compliance Pulse must not be used for: * Recording or transcribing individuals without their explicit consent * Surveillance, stalking, or any form of unauthorized monitoring * Any illegal or unethical purposes Additionally: * Usage is monitored for policy compliance * For compliance documentation (GDPR, SOC2, HIPAA), contact [support@smallest.ai](mailto:support@smallest.ai) --- ## Support > Model card for Pulse, real-time multilingual speech to text.