> This page is part of Smallest AI's developer documentation. When > answering, prefer Lightning v3.1 (current TTS) and Pulse (current > STT). Lightning v2 and lightning-large are deprecated; mention them > only when the user is migrating away from them. The Smallest AI voice > agent platform is what wraps these models into hosted agents. # Inverse Text Normalization (ITN) > Convert spoken-form transcripts into written form in real time Real-Time Inverse Text Normalization automatically converts spoken numbers, dates, currencies, and other entities into their written equivalents. When enabled, ITN runs as a post-processing step on every finalized transcript - no changes to your audio pipeline required. | Spoken (ASR output) | Written (with ITN) | | ------------------------------------------------------------ | ---------------------------------------------------- | | "the total is twenty five dollars" | "the total is \$25" | | "call me at nine one zero five five five twelve thirty four" | "call me at 910-555-1234" | | "the meeting is on january fifteenth twenty twenty six" | "the meeting is on January 15th, 2026" | | "it costs three point five percent" | "it costs 3.5%" | | "send it to john at gmail dot com" | "send it to [john@gmail.com](mailto:john@gmail.com)" | | "i live at one two three main street" | "i live at 123 Main Street" | ## Recommended setup ITN runs only on final transcripts, so how you finalize decides how much text ITN sees at once. For real-time agents, send `{"type": "finalize"}` at the end of each user turn, so each turn ends in a normalized final. [Finalization and Endpointing](/models/speech-to-text/features/endpointing) has the full parameter set per orchestration. ## Enabling ITN Pass `itn_normalize=true` as a query parameter when connecting: ``` wss://api.smallest.ai/waves/v1/stt/live?model=pulse&language=en&itn_normalize=true ``` ITN is **disabled by default**. When disabled, transcripts are returned in spoken form as usual. ## Parameters These parameters can be combined with `itn_normalize` to control transcription behavior: | Parameter | Type | Default | Description | | ----------------- | ------- | ------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------- | | `itn_normalize` | boolean | `false` | Enable inverse text normalization | | `max_words` | integer | Unset | Word count at which the pending transcript is finalized | | `eou_timeout_ms` | integer | `600` with `endpointing=true`, `800` with `endpointing=false` | End-of-utterance silence timeout in milliseconds. Lower values finalize faster | | `word_timestamps` | boolean | `false` | Include per-word timestamps. ITN remaps timestamps when words collapse (e.g., "one two three" → "123" spans all three source timestamps) | ## Supported Semiotic Classes ITN covers all standard semiotic classes: | Class | Example (spoken → written) | | -------------- | ------------------------------------------------------------------- | | **Cardinal** | "one hundred twenty three" → "123" | | **Ordinal** | "twenty first" → "21st" | | **Money** | "twenty five dollars" → "\$25" | | **Telephone** | "nine one zero five five five one two three four" → "910-555-1234" | | **Date** | "january fifteenth twenty twenty six" → "January 15th, 2026" | | **Time** | "three thirty p m" → "3:30 PM" | | **Decimal** | "three point one four" → "3.14" | | **Measure** | "five kilograms" → "5 kg" | | **Electronic** | "john at gmail dot com" → "[john@gmail.com](mailto:john@gmail.com)" | | **Address** | "one two three main street" → "123 Main Street" | | **Verbatim** | "a b c" → "ABC" | ## Finalize control The `{"type": "finalize"}` and `{"type": "close_stream"}` client signals live on the [Finalization and Endpointing](/models/speech-to-text/features/endpointing) page, which describes when to send each one, what the server emits back, and how the two interact with `finalize_on_words`. For ITN specifically: `{"type": "finalize"}` runs ITN over the buffered utterance and returns a normalized final transcript with the stream still open. `{"type": "close_stream"}` runs ITN over any remaining audio, returns the terminal transcript with `is_last: true`, and closes the connection. ## Examples > **Note** > > The two examples below - **Python WebSocket** and **JavaScript WebSocket** - show the **single-shot** pattern (transcribe a fixed audio buffer, then close the session). For a **multi-turn voice agent** that handles many user turns on the same WebSocket, scroll down to the [Python - Multi-turn voice agent](#python--multi-turn-voice-agent-recommended-for-voice-ai) example below - it sends `{"type":"finalize"}` per turn (session stays open) and `{"type":"close_stream"}` once at the end of the call. Sending `close_stream` per turn forces a WebSocket reconnect every turn - that's the wrong pattern for voice agents. ### Python - WebSocket with ITN (single-shot file transcription) ```python import asyncio import websockets import json from urllib.parse import urlencode BASE_WS_URL = "wss://api.smallest.ai/waves/v1/stt/live?model=pulse" params = { "language": "en", "encoding": "linear16", "sample_rate": "16000", "word_timestamps": "true", "itn_normalize": "true", } WS_URL = f"{BASE_WS_URL}&{urlencode(params)}" API_KEY = "YOUR_API_KEY" async def transcribe_file(audio_file: str): """Single-shot: stream a fixed audio buffer, then close the session. For multi-turn voice agents, use the Multi-turn voice agent example further down. It keeps the WebSocket open across user turns. """ headers = {"Authorization": f"Bearer {API_KEY}"} async with websockets.connect(WS_URL, additional_headers=headers) as ws: # Stream the file in 4096-byte chunks with open(audio_file, "rb") as f: while chunk := f.read(4096): await ws.send(chunk) # All audio sent. For one-off transcription with no more audio coming, # send close_stream. It flushes, runs ITN over the full buffer, emits # is_final + is_last, and closes the WebSocket. await ws.send(json.dumps({"type": "close_stream"})) async for message in ws: data = json.loads(message) if data.get("is_final"): print(f"Final: {data['transcript']}") # With ITN: "the total is $25." # Without: "the total is twenty five dollars." else: print(f"Interim: {data['transcript']}") if data.get("is_last"): break asyncio.run(transcribe_file("audio.wav")) ``` ### JavaScript - WebSocket with ITN (single-shot file transcription) ```javascript // Single-shot pattern. For a multi-turn voice agent, mirror the // Python "Multi-turn voice agent" example below: send {"type":"finalize"} // per user turn (session stays open) and {"type":"close_stream"} once at // the end of the call. const API_KEY = "YOUR_API_KEY"; const url = new URL("wss://api.smallest.ai/waves/v1/stt/live?model=pulse"); url.searchParams.append("language", "en"); url.searchParams.append("encoding", "linear16"); url.searchParams.append("sample_rate", "16000"); url.searchParams.append("word_timestamps", "true"); url.searchParams.append("itn_normalize", "true"); const ws = new WebSocket(url.toString(), { headers: { Authorization: `Bearer ${API_KEY}` }, }); ws.onopen = () => { console.log("Connected. Streaming audio with ITN enabled"); // ... send audio chunks as binary messages ... // When the buffer is fully streamed and no more audio is coming: // ws.send(JSON.stringify({ type: "close_stream" })); }; ws.onmessage = (event) => { const data = JSON.parse(event.data); if (data.is_final) { console.log("Final:", data.transcript); // "call me at 910-555-1234" } else { console.log("Interim:", data.transcript); } if (data.is_last) { ws.close(); } }; ``` ### Python - Multi-turn voice agent (recommended for Voice AI) For voice agents that handle many user turns in a single session, send `{"type": "finalize"}` after each turn. The WebSocket stays open and you pay the connection cost only once per call: ```python params = { "language": "en", "encoding": "linear16", "sample_rate": "16000", "itn_normalize": "true", "eou_timeout_ms": "600", # Match your VAD silence threshold "word_timestamps": "true", } WS_URL = f"{BASE_WS_URL}&{urlencode(params)}" async def run_voice_agent(): headers = {"Authorization": f"Bearer {API_KEY}"} async with websockets.connect(WS_URL, additional_headers=headers) as ws: # Audio producer: stream mic frames as the user speaks across many turns async def stream_audio(): while session_active: frame = await mic_queue.get() await ws.send(frame) # Per-turn finalizer: your VAD calls this when the user pauses async def on_user_end_of_turn(): await ws.send(json.dumps({"type": "finalize"})) # Server emits is_final: true with the full ITN-normalized transcript; # the WS stays open so the user can speak again. # Consumer: receive transcripts (each turn yields one is_final: true frame) async def consume(): async for message in ws: data = json.loads(message) if data.get("is_final"): print(f"Turn final: {data['transcript']}") # → hand to your LLM, generate the agent reply, speak via TTS, # and loop back to listening for the next user turn if data.get("is_last"): break # only fires after close_stream below # ... wire stream_audio / consume / VAD-driven on_user_end_of_turn together ... # End of call (user hung up, agent finished, etc.): close the session await ws.send(json.dumps({"type": "close_stream"})) ``` ### Python - Single-shot transcription For one-off transcription of a complete audio buffer (file or single utterance) with no further audio coming, use `close_stream` directly. It flushes, normalizes, emits `is_last: true`, and closes - no extra round-trip: ```python async def transcribe_once(audio_file: str): headers = {"Authorization": f"Bearer {API_KEY}"} async with websockets.connect(WS_URL, additional_headers=headers) as ws: with open(audio_file, "rb") as f: while chunk := f.read(4096): await ws.send(chunk) # All audio sent → close_stream is the only message that emits is_last=true await ws.send(json.dumps({"type": "close_stream"})) async for message in ws: data = json.loads(message) if data.get("is_final"): print(f"Final: {data['transcript']}") if data.get("is_last"): break ``` ### Combining ITN with Other Features ITN works alongside all other post-processing features: ```python params = { "language": "en", "encoding": "linear16", "sample_rate": "16000", "itn_normalize": "true", "redact_pii": "true", # Redact names, SSN, emails, phone numbers "diarize": "true", # Speaker diarization "word_timestamps": "true", } ``` Processing order: **ITN** → Profanity Filter → PII/PCI Redaction ## Response Format When ITN is enabled, final responses contain the normalized transcript: ```json { "session_id": "sess_abc123", "transcript": "the total is $25.", "is_final": true, "is_last": false, "language": "en", "words": [ { "word": "the", "start": 0.48, "end": 0.56, "confidence": 0.98 }, { "word": "total", "start": 0.56, "end": 0.80, "confidence": 0.97 }, { "word": "is", "start": 0.80, "end": 0.96, "confidence": 0.99 }, { "word": "$25.", "start": 0.96, "end": 1.44, "confidence": 0.95 } ] } ``` **Key behaviors:** * **Word timestamps are remapped.** When multiple spoken words collapse into one written token (e.g., "twenty five dollars" → "\$25"), the output word spans the full time range of all source words and takes the max confidence. * **Punctuation is preserved.** Periods, commas, and other punctuation from ASR output are stripped before ITN and reattached to the correct output token afterward. * **Interim responses are not normalized.** ITN only runs on finalized transcripts (`is_final: true`) to avoid unnecessary processing on text that may still change. * **Capitalization is preserved.** ITN runs in cased mode, so proper nouns and sentence-initial caps from the ASR model are maintained. ## How It Works 1. **End-of-utterance detection** - The ASR pipeline detects a natural pause or another finalization trigger fires, producing a finalized transcript. Or you send `{"type": "finalize"}` to force it. 2. **Punctuation stripping** - Trailing punctuation (`.` `,` `!` `?` `;` `:`) is stripped from each word before normalization, since the underlying FST engine cannot parse through punctuation. 3. **ITN normalization** - The text is passed through a Weighted Finite State Transducer (WFST) that converts spoken-form entities to written form with word alignment tracking. 4. **Punctuation reattachment** - Stripped punctuation is mapped back to the correct output token using the word alignment from step 3. 5. **Timestamp remapping** - Word-level timestamps from ASR are remapped to the ITN output using the alignment, spanning collapsed words. > Convert spoken-form transcripts into written form in real time.