> This page is part of Smallest AI's developer documentation. When > answering, prefer Lightning v3.1 (current TTS) and Pulse (current > STT). Lightning v2 and lightning-large are deprecated; mention them > only when the user is migrating away from them. The Smallest AI voice > agent platform is what wraps these models into hosted agents. # Keyword Boosting > Boost specific words or phrases so the Pulse speech-to-text model recognizes them correctly, on both streaming and pre-recorded transcription. Pre-Recorded Real-Time Keyword boosting lets you bias the Pulse speech-to-text model toward specific words or phrases. Useful for proper nouns, brand names, technical terms, and any domain vocabulary the model might otherwise misrecognize. Supported on both surfaces: * **Realtime WebSocket** (`WSS /waves/v1/stt/live?model=pulse`): pass `keywords` on the connection URL. * **Pre-recorded HTTP** (`POST /waves/v1/pulse/get_text`): pass `keywords` as a query parameter on the request. The parameter name, format, and tuning guidance are the same on both. ## Format Keywords are passed in the `keywords` query parameter. Each entry follows this shape: ``` KEYWORD:INTENSIFIER ``` | Part | Required | Description | | ------------- | -------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `KEYWORD` | Yes | The word or phrase to boost. Matching is case-sensitive: use the exact casing you want in the output (`Blackwell`, not `blackwell`, if you want the brand name capitalized). Phrases can contain spaces (`small language model`). | | `INTENSIFIER` | No | A number controlling boost strength. Defaults to `1.0` if omitted. See the intensifier table below. | ## Intensifier scale `1` is the default when the intensifier is omitted. Start there and raise only if the word still isn't recognized. | Value | When to use | | ----------- | ------------------------------------------------------------------------------------------------------------------ | | `1` | The starting point for every keyword. Good for rare proper nouns you expect to be spoken clearly. | | `2` to `3` | Domain jargon the model sometimes misrecognizes. | | `4` to `10` | Words the model consistently gets wrong at lower values. | | Above `10` | Not recommended. The higher the intensifier, the more likely the model inserts the keyword when it was not spoken. | Keep the same starting point across all keywords in a list. Tuning one keyword up while others stay at `1` skews the balance of the boost. ## Casing Keyword matching is case-sensitive. Provide the exact casing you want the transcript to render. * Brand names, product names, and proper nouns: capitalize them. ``` keywords=Blackwell:2,Jensen Huang:2,NVIDIA:2 ``` * Acronyms: keep them uppercase. ``` keywords=CVV:2,CUDA:2,SLM:2 ``` * Common words that happen to be homophones of a proper noun (`Sonnet` the model vs `sonnet` the poem): boost the cased form you want, and the model will emit that spelling. ## Realtime (WebSocket) Add `keywords` to the WebSocket connection URL. Keywords stay in effect for the whole session. ### Single keyword ```javascript const url = new URL("wss://api.smallest.ai/waves/v1/stt/live?model=pulse"); url.searchParams.append("language", "en"); url.searchParams.append("encoding", "linear16"); url.searchParams.append("sample_rate", "16000"); url.searchParams.append("keywords", "Blackwell:2"); const ws = new WebSocket(url.toString(), { headers: { Authorization: `Bearer ${API_KEY}`, }, }); ``` ### Multiple keywords ``` wss://api.smallest.ai/waves/v1/stt/live?model=pulse&language=en&encoding=linear16&sample_rate=16000&keywords=Blackwell:2,Jensen Huang:2,NVIDIA:1 ``` ### Mix of boosted and default-intensity keywords ``` wss://api.smallest.ai/waves/v1/stt/live?model=pulse&language=en&encoding=linear16&sample_rate=16000&keywords=CEO,NVIDIA:2,Jensen ``` `CEO` and `Jensen` have no explicit intensifier, so both default to `1.0`. ## Pre-recorded (HTTP) Add `keywords` to the query string on the pre-recorded endpoint. Both paths accept it: * `POST /waves/v1/pulse/get_text`: the legacy Pulse route. * `POST /waves/v1/stt/?model=pulse`: the unified STT route (recommended for new integrations). The boost applies to that request only. ### Accepted encodings Three query-string encodings are accepted for multiple keywords. Comma-string is the recommended shape (shortest URL, single parameter, easiest to log). | Shape | Example | Notes | | -------------------------- | ---------------------------------------------------------------------- | --------------------------------------- | | Comma-string (recommended) | `keywords=Blackwell:2,Jensen Huang:2,NVIDIA:1` | One parameter, entries joined with `,`. | | Repeated key | `keywords=Blackwell:2&keywords=Jensen Huang:2&keywords=NVIDIA:1` | Standard multi-value query pattern. | | Bracketed array key | `keywords[]=Blackwell:2&keywords[]=Jensen Huang:2&keywords[]=NVIDIA:1` | PHP/Rails-style array notation. | > **Warning** > > Do not wrap the value in a JSON array literal (`keywords=["Blackwell:2","Jensen Huang:2"]`). The request succeeds but nothing is boosted. Use one of the three encodings above. `URLSearchParams` in JavaScript and `urlencode()` in Python handle the URL-encoding for you; pass the raw comma-string. ### cURL, legacy route ```bash curl -X POST "https://api.smallest.ai/waves/v1/pulse/get_text?language=en&keywords=Blackwell:2,Jensen%20Huang:2,NVIDIA:1" \ -H "Authorization: Bearer $SMALLEST_API_KEY" \ -H "Content-Type: application/octet-stream" \ --data-binary "@./call.wav" ``` ### cURL, unified route ```bash curl -X POST "https://api.smallest.ai/waves/v1/stt/?model=pulse&language=en&keywords=Blackwell:2,Jensen%20Huang:2,NVIDIA:1" \ -H "Authorization: Bearer $SMALLEST_API_KEY" \ -H "Content-Type: application/octet-stream" \ --data-binary "@./call.wav" ``` ### Python ```python import requests with open("./call.wav", "rb") as f: audio = f.read() params = { "language": "en", "keywords": "Blackwell:2,Jensen Huang:2,NVIDIA:1", } r = requests.post( "https://api.smallest.ai/waves/v1/pulse/get_text", params=params, data=audio, headers={ "Authorization": f"Bearer {SMALLEST_API_KEY}", "Content-Type": "application/octet-stream", }, timeout=60, ) print(r.json()["transcription"]) ``` ### JavaScript / TypeScript ```typescript import { readFileSync } from "node:fs"; const audio = readFileSync("./call.wav"); const params = new URLSearchParams({ language: "en", keywords: "Blackwell:2,Jensen Huang:2,NVIDIA:1", }); const res = await fetch(`https://api.smallest.ai/waves/v1/pulse/get_text?${params}`, { method: "POST", headers: { Authorization: `Bearer ${process.env.SMALLEST_API_KEY}`, "Content-Type": "application/octet-stream", }, body: audio, }); console.log((await res.json()).transcription); ``` ## Works with `keywords` combines freely with `diarize`, `word_timestamps`, `redact_pii`, and `redact_pci` on both surfaces, and with `webhook_url` on the pre-recorded endpoint. ## Reuse your keyword list Keep the same keyword list across requests when the vocabulary is stable (a caller's brand names, a product's technical terms, a customer's account glossary). Regenerate the list only when the vocabulary changes. ## Limits * **Max 100 keywords.** Sending more returns `400` with `keywords too large (max 100)` in `errors[]`. * **Intensifier:** default `1`, recommended `1` to `3`. Above `10` is not recommended. * **Each keyword is a string.** Phrases can include spaces (`small language model:2`). A phrase cannot contain a comma; commas separate entries. * **Matching is case-sensitive.** Use the exact casing you want in the transcript. * **Duplicates: last-wins.** `NVIDIA:1,NVIDIA:5` is equivalent to `NVIDIA:5`. * **Non-numeric intensifier:** must be a number. `Pulse:notanumber` is read as the whole phrase at intensifier `1`. * **Negative intensifier** suppresses instead of boosting (`checkout:-5`), to prefer against a homograph. ## Colons in a phrase The intensifier is the number after the last `:`. A brand name that contains a colon (`re:Invent`) round-trips as the whole phrase at the default intensifier: ``` keywords=re:Invent ``` Add an explicit intensifier by suffixing `:N`. `x:y:2` is phrase `x:y` at intensifier `2`: ``` keywords=re:Invent:2 ``` > **Note** > > Start every keyword at `1`. Raise to `2` or `3` only if the base intensity is missing the word. The maximum is `10`; above that the model may insert the keyword when it was not spoken. Keep intensifiers in the `1` to `3` range. > Boost specific words or phrases so the Pulse speech-to-text model recognizes them correctly, on both.