Inverse Text Normalization (ITN)
Inverse Text Normalization automatically converts spoken numbers, dates, currencies, and other entities into their written equivalents. When enabled, ITN runs as a post-processing step on every finalized transcript - no changes to your audio pipeline required.
Recommended setup
ITN runs only on final transcripts, so how you finalize decides how much text ITN sees at once. For real-time agents, send {"type": "finalize"} at the end of each user turn, so each turn ends in a normalized final. Finalization and Endpointing has the full parameter set per orchestration.
Enabling ITN
Pass itn_normalize=true as a query parameter when connecting:
ITN is disabled by default. When disabled, transcripts are returned in spoken form as usual.
Parameters
These parameters can be combined with itn_normalize to control transcription behavior:
Supported Semiotic Classes
ITN covers all standard semiotic classes:
Finalize control
The {"type": "finalize"} and {"type": "close_stream"} client signals live on the Finalization and Endpointing page, which describes when to send each one, what the server emits back, and how the two interact with finalize_on_words.
For ITN specifically: {"type": "finalize"} runs ITN over the buffered utterance and returns a normalized final transcript with the stream still open. {"type": "close_stream"} runs ITN over any remaining audio, returns the terminal transcript with is_last: true, and closes the connection.
Examples
The two examples below - Python WebSocket and JavaScript WebSocket - show the single-shot pattern (transcribe a fixed audio buffer, then close the session). For a multi-turn voice agent that handles many user turns on the same WebSocket, scroll down to the Python - Multi-turn voice agent example below - it sends {"type":"finalize"} per turn (session stays open) and {"type":"close_stream"} once at the end of the call. Sending close_stream per turn forces a WebSocket reconnect every turn - that’s the wrong pattern for voice agents.
Python - WebSocket with ITN (single-shot file transcription)
JavaScript - WebSocket with ITN (single-shot file transcription)
Python - Multi-turn voice agent (recommended for Voice AI)
For voice agents that handle many user turns in a single session, send {"type": "finalize"} after each turn. The WebSocket stays open and you pay the connection cost only once per call:
Python - Single-shot transcription
For one-off transcription of a complete audio buffer (file or single utterance) with no further audio coming, use close_stream directly. It flushes, normalizes, emits is_last: true, and closes - no extra round-trip:
Combining ITN with Other Features
ITN works alongside all other post-processing features:
Processing order: ITN → Profanity Filter → PII/PCI Redaction
Response Format
When ITN is enabled, final responses contain the normalized transcript:
Key behaviors:
- Word timestamps are remapped. When multiple spoken words collapse into one written token (e.g., “twenty five dollars” → “$25”), the output word spans the full time range of all source words and takes the max confidence.
- Punctuation is preserved. Periods, commas, and other punctuation from ASR output are stripped before ITN and reattached to the correct output token afterward.
- Interim responses are not normalized. ITN only runs on finalized transcripts (
is_final: true) to avoid unnecessary processing on text that may still change. - Capitalization is preserved. ITN runs in cased mode, so proper nouns and sentence-initial caps from the ASR model are maintained.
How It Works
- End-of-utterance detection - The ASR pipeline detects a natural pause or another finalization trigger fires, producing a finalized transcript. Or you send
{"type": "finalize"}to force it. - Punctuation stripping - Trailing punctuation (
.,!?;:) is stripped from each word before normalization, since the underlying FST engine cannot parse through punctuation. - ITN normalization - The text is passed through a Weighted Finite State Transducer (WFST) that converts spoken-form entities to written form with word alignment tracking.
- Punctuation reattachment - Stripped punctuation is mapped back to the correct output token using the word alignment from step 3.
- Timestamp remapping - Word-level timestamps from ASR are remapped to the ITN output using the alignment, spanning collapsed words.