East Asian languages
Pulse streaming accuracy on Mandarin, Cantonese, Japanese and Korean.
East Asian languages - Multi-dataset (Streaming)
WER for the four East Asian languages on the streaming endpoint (US region). Three datasets per language covering read speech (FLEURS), conversational/crowdsourced speech (Common Voice 25), and language-specific corpora (JSUT, Zeroth-Korean, MDCC, AISHELL-1). Compared head-to-head against Deepgram Nova-3. Lower WER is better.
Pulse averages 10.91% WER vs Deepgram Nova-3’s 15.50% across the four East Asian languages - Pulse leads on 10 of 12 dataset rows, with the largest gains on Japanese CV-25 and Cantonese CV-25.
These four languages stream from wss://api.us.smallest.ai/waves/v1/stt/live?model=pulse only (US region). See the Pulse model card for the region-routing details.