Features

View as Markdown

The Real-Time Pulse STT WebSocket API supports the following features:

Available Features

Word Timestamps

Get precise timing information for each word in the transcription with confidence scores

Language Detection

Automatically detect the language of the audio

Sentence Timestamps (Utterances)

Get sentence-level transcription segments with timing information

PII & PCI Redaction

Automatically redact personally identifiable information and payment card information

Speaker Diarization

Identify and label different speakers in the audio with speaker confidence scores

Keyword Boosting

Boost recognition accuracy for specific words, brand names, and domain terms

Punctuation Formatting

Control punctuation and capitalization formatting in transcripts

End-of-Utterance Timeout

Control how long Pulse waits after speech ends before finalizing the transcript

Inverse Text Normalization

Convert spoken-form numbers, dates, and currencies into written form

Finalize Control

Take manual control of when transcripts are finalized using finalize_on_words and max_words

VAD Events

Emit acoustic speech_started / speech_ended frames interleaved with the transcription stream. Independent of transcript finalization.

Endpointing

Tune the end-of-utterance silence timeout with eou_timeout_ms for interactive vs dictation-style workloads.

Age & Gender Detection

Optional per-speaker age and gender inference alongside the transcript.

Emotion Detection

Optional per-utterance emotion classification alongside the transcript.