Skip to navigation

Inverse Text Normalization (ITN)

Convert spoken-form transcripts into written form in real time.
View as Markdown
Real-Time

Inverse Text Normalization automatically converts spoken numbers, dates, currencies, and other entities into their written equivalents. When enabled, ITN runs as a post-processing step on every finalized transcript - no changes to your audio pipeline required.

Spoken (ASR output)Written (with ITN)
“the total is twenty five dollars""the total is $25"
"call me at nine one zero five five five twelve thirty four""call me at 910-555-1234"
"the meeting is on january fifteenth twenty twenty six""the meeting is on January 15th, 2026"
"it costs three point five percent""it costs 3.5%"
"send it to john at gmail dot com""send it to john@gmail.com"
"i live at one two three main street""i live at 123 Main Street”

ITN runs only on final transcripts, so how you finalize decides how much text ITN sees at once. For real-time agents, send {"type": "finalize"} at the end of each user turn, so each turn ends in a normalized final. Finalization and Endpointing has the full parameter set per orchestration.

Enabling ITN

Pass itn_normalize=true as a query parameter when connecting:

wss://api.smallest.ai/waves/v1/stt/live?model=pulse&language=en&itn_normalize=true

ITN is disabled by default. When disabled, transcripts are returned in spoken form as usual.

Parameters

These parameters can be combined with itn_normalize to control transcription behavior:

ParameterTypeDefaultDescription
itn_normalizebooleanfalseEnable inverse text normalization
max_wordsintegerUnsetWord count at which the pending transcript is finalized
eou_timeout_msinteger600 with endpointing=true, 800 with endpointing=falseEnd-of-utterance silence timeout in milliseconds. Lower values finalize faster
word_timestampsbooleanfalseInclude per-word timestamps. ITN remaps timestamps when words collapse (e.g., “one two three” → “123” spans all three source timestamps)

Supported Semiotic Classes

ITN covers all standard semiotic classes:

ClassExample (spoken → written)
Cardinal”one hundred twenty three” → “123”
Ordinal”twenty first” → “21st”
Money”twenty five dollars” → “$25”
Telephone”nine one zero five five five one two three four” → “910-555-1234”
Date”january fifteenth twenty twenty six” → “January 15th, 2026”
Time”three thirty p m” → “3:30 PM”
Decimal”three point one four” → “3.14”
Measure”five kilograms” → “5 kg”
Electronic”john at gmail dot com” → “john@gmail.com”
Address”one two three main street” → “123 Main Street”
Verbatim”a b c” → “ABC”

Finalize control

The {"type": "finalize"} and {"type": "close_stream"} client signals live on the Finalization and Endpointing page, which describes when to send each one, what the server emits back, and how the two interact with finalize_on_words.

For ITN specifically: {"type": "finalize"} runs ITN over the buffered utterance and returns a normalized final transcript with the stream still open. {"type": "close_stream"} runs ITN over any remaining audio, returns the terminal transcript with is_last: true, and closes the connection.

Examples

The two examples below - Python WebSocket and JavaScript WebSocket - show the single-shot pattern (transcribe a fixed audio buffer, then close the session). For a multi-turn voice agent that handles many user turns on the same WebSocket, scroll down to the Python - Multi-turn voice agent example below - it sends {"type":"finalize"} per turn (session stays open) and {"type":"close_stream"} once at the end of the call. Sending close_stream per turn forces a WebSocket reconnect every turn - that’s the wrong pattern for voice agents.

Python - WebSocket with ITN (single-shot file transcription)

import asyncio
import websockets
import json
from urllib.parse import urlencode
BASE_WS_URL = "wss://api.smallest.ai/waves/v1/stt/live?model=pulse"
params = {
"language": "en",
"encoding": "linear16",
"sample_rate": "16000",
"word_timestamps": "true",
"itn_normalize": "true",
}
WS_URL = f"{BASE_WS_URL}&{urlencode(params)}"
API_KEY = "YOUR_API_KEY"
async def transcribe_file(audio_file: str):
"""Single-shot: stream a fixed audio buffer, then close the session.
For multi-turn voice agents, use the Multi-turn voice agent example
further down. It keeps the WebSocket open across user turns.
"""
headers = {"Authorization": f"Bearer {API_KEY}"}
async with websockets.connect(WS_URL, additional_headers=headers) as ws:
# Stream the file in 4096-byte chunks
with open(audio_file, "rb") as f:
while chunk := f.read(4096):
await ws.send(chunk)
# All audio sent. For one-off transcription with no more audio coming,
# send close_stream. It flushes, runs ITN over the full buffer, emits
# is_final + is_last, and closes the WebSocket.
await ws.send(json.dumps({"type": "close_stream"}))
async for message in ws:
data = json.loads(message)
if data.get("is_final"):
print(f"Final: {data['transcript']}")
# With ITN: "the total is $25."
# Without: "the total is twenty five dollars."
else:
print(f"Interim: {data['transcript']}")
if data.get("is_last"):
break
asyncio.run(transcribe_file("audio.wav"))

JavaScript - WebSocket with ITN (single-shot file transcription)

// Single-shot pattern. For a multi-turn voice agent, mirror the
// Python "Multi-turn voice agent" example below: send {"type":"finalize"}
// per user turn (session stays open) and {"type":"close_stream"} once at
// the end of the call.
const API_KEY = "YOUR_API_KEY";
const url = new URL("wss://api.smallest.ai/waves/v1/stt/live?model=pulse");
url.searchParams.append("language", "en");
url.searchParams.append("encoding", "linear16");
url.searchParams.append("sample_rate", "16000");
url.searchParams.append("word_timestamps", "true");
url.searchParams.append("itn_normalize", "true");
const ws = new WebSocket(url.toString(), {
headers: { Authorization: `Bearer ${API_KEY}` },
});
ws.onopen = () => {
console.log("Connected. Streaming audio with ITN enabled");
// ... send audio chunks as binary messages ...
// When the buffer is fully streamed and no more audio is coming:
// ws.send(JSON.stringify({ type: "close_stream" }));
};
ws.onmessage = (event) => {
const data = JSON.parse(event.data);
if (data.is_final) {
console.log("Final:", data.transcript);
// "call me at 910-555-1234"
} else {
console.log("Interim:", data.transcript);
}
if (data.is_last) {
ws.close();
}
};

For voice agents that handle many user turns in a single session, send {"type": "finalize"} after each turn. The WebSocket stays open and you pay the connection cost only once per call:

params = {
"language": "en",
"encoding": "linear16",
"sample_rate": "16000",
"itn_normalize": "true",
"eou_timeout_ms": "600", # Match your VAD silence threshold
"word_timestamps": "true",
}
WS_URL = f"{BASE_WS_URL}&{urlencode(params)}"
async def run_voice_agent():
headers = {"Authorization": f"Bearer {API_KEY}"}
async with websockets.connect(WS_URL, additional_headers=headers) as ws:
# Audio producer: stream mic frames as the user speaks across many turns
async def stream_audio():
while session_active:
frame = await mic_queue.get()
await ws.send(frame)
# Per-turn finalizer: your VAD calls this when the user pauses
async def on_user_end_of_turn():
await ws.send(json.dumps({"type": "finalize"}))
# Server emits is_final: true with the full ITN-normalized transcript;
# the WS stays open so the user can speak again.
# Consumer: receive transcripts (each turn yields one is_final: true frame)
async def consume():
async for message in ws:
data = json.loads(message)
if data.get("is_final"):
print(f"Turn final: {data['transcript']}")
# → hand to your LLM, generate the agent reply, speak via TTS,
# and loop back to listening for the next user turn
if data.get("is_last"):
break # only fires after close_stream below
# ... wire stream_audio / consume / VAD-driven on_user_end_of_turn together ...
# End of call (user hung up, agent finished, etc.): close the session
await ws.send(json.dumps({"type": "close_stream"}))

Python - Single-shot transcription

For one-off transcription of a complete audio buffer (file or single utterance) with no further audio coming, use close_stream directly. It flushes, normalizes, emits is_last: true, and closes - no extra round-trip:

async def transcribe_once(audio_file: str):
headers = {"Authorization": f"Bearer {API_KEY}"}
async with websockets.connect(WS_URL, additional_headers=headers) as ws:
with open(audio_file, "rb") as f:
while chunk := f.read(4096):
await ws.send(chunk)
# All audio sent → close_stream is the only message that emits is_last=true
await ws.send(json.dumps({"type": "close_stream"}))
async for message in ws:
data = json.loads(message)
if data.get("is_final"):
print(f"Final: {data['transcript']}")
if data.get("is_last"):
break

Combining ITN with Other Features

ITN works alongside all other post-processing features:

params = {
"language": "en",
"encoding": "linear16",
"sample_rate": "16000",
"itn_normalize": "true",
"redact_pii": "true", # Redact names, SSN, emails, phone numbers
"diarize": "true", # Speaker diarization
"word_timestamps": "true",
}

Processing order: ITN → Profanity Filter → PII/PCI Redaction

Response Format

When ITN is enabled, final responses contain the normalized transcript:

{
"session_id": "sess_abc123",
"transcript": "the total is $25.",
"is_final": true,
"is_last": false,
"language": "en",
"words": [
{ "word": "the", "start": 0.48, "end": 0.56, "confidence": 0.98 },
{ "word": "total", "start": 0.56, "end": 0.80, "confidence": 0.97 },
{ "word": "is", "start": 0.80, "end": 0.96, "confidence": 0.99 },
{ "word": "$25.", "start": 0.96, "end": 1.44, "confidence": 0.95 }
]
}

Key behaviors:

  • Word timestamps are remapped. When multiple spoken words collapse into one written token (e.g., “twenty five dollars” → “$25”), the output word spans the full time range of all source words and takes the max confidence.
  • Punctuation is preserved. Periods, commas, and other punctuation from ASR output are stripped before ITN and reattached to the correct output token afterward.
  • Interim responses are not normalized. ITN only runs on finalized transcripts (is_final: true) to avoid unnecessary processing on text that may still change.
  • Capitalization is preserved. ITN runs in cased mode, so proper nouns and sentence-initial caps from the ASR model are maintained.

How It Works

  1. End-of-utterance detection - The ASR pipeline detects a natural pause or another finalization trigger fires, producing a finalized transcript. Or you send {"type": "finalize"} to force it.
  2. Punctuation stripping - Trailing punctuation (. , ! ? ; :) is stripped from each word before normalization, since the underlying FST engine cannot parse through punctuation.
  3. ITN normalization - The text is passed through a Weighted Finite State Transducer (WFST) that converts spoken-form entities to written form with word alignment tracking.
  4. Punctuation reattachment - Stripped punctuation is mapped back to the correct output token using the word alignment from step 3.
  5. Timestamp remapping - Word-level timestamps from ASR are remapped to the ITN output using the alignment, spanning collapsed words.