Skip to navigation

Transcribe (Pre-recorded)

View as Markdown

Transcribe an audio file. The model is chosen via ?model=:

  • ?model=pulse-pro: English-only, leaderboard-ranked accuracy. Raw bytes only; pass webhook_url to receive transcription asynchronously on long files.
  • ?model=pulse: multilingual transcription (21 streaming + 12 pre-recorded languages), supports both raw bytes and audio-by-URL.

When to use this

Use this endpoint when you have a complete audio file (call recording, voicemail, podcast episode) and want the transcript back in one response. For live transcription as audio arrives, use the realtime WebSocket endpoint (WS /waves/v1/stt/live?model=pulse).

Pulse Pro is HTTP-only.

Input methods

  • Raw bytes: Content-Type: application/octet-stream with the audio in the body. All knobs are query parameters.
  • URL (?model=pulse only): Content-Type: application/json with {"url": "..."} in the body.

Examples

cURL: Pulse Pro, sync

curl -X POST "https://api.smallest.ai/waves/v1/stt/?model=pulse-pro&language=en&word_timestamps=true" \
-H "Authorization: Bearer $SMALLEST_API_KEY" \
-H "Content-Type: application/octet-stream" \
--data-binary "@./call.wav"

cURL: Pulse Pro, async via webhook

curl -X POST "https://api.smallest.ai/waves/v1/stt/?model=pulse-pro&language=en&webhook_url=https://your.app/cb" \
-H "Authorization: Bearer $SMALLEST_API_KEY" \
-H "Content-Type: application/octet-stream" \
--data-binary "@./call.wav"

Returns 200 { "status": "processing", "request_id": "..." } immediately. The webhook receives the full transcription when ready.

cURL: Pulse, audio-by-URL

curl -X POST "https://api.smallest.ai/waves/v1/stt/?model=pulse&language=en" \
-H "Authorization: Bearer $SMALLEST_API_KEY" \
-H "Content-Type: application/json" \
-d '{"url": "https://your-bucket.s3.amazonaws.com/call.wav"}'

Python

import requests
with open("./call.wav", "rb") as f:
audio = f.read()
r = requests.post(
"https://api.smallest.ai/waves/v1/stt/",
params={"model": "pulse-pro", "language": "en", "word_timestamps": "true"},
headers={"Authorization": f"Bearer {API_KEY}", "Content-Type": "application/octet-stream"},
data=audio,
)
r.raise_for_status()
print(r.json()["transcription"])

JavaScript / TypeScript

import { readFileSync } from "node:fs";
const audio = readFileSync("./call.wav");
const params = new URLSearchParams({ model: "pulse-pro", language: "en", word_timestamps: "true" });
const res = await fetch(`https://api.smallest.ai/waves/v1/stt/?${params}`, {
method: "POST",
headers: { Authorization: `Bearer ${process.env.SMALLEST_API_KEY}`, "Content-Type": "application/octet-stream" },
body: audio,
});
console.log((await res.json()).transcription);

Common gotchas

  • model is required. Missing or invalid values return 400 with an enum-validation error.
  • Pulse Pro is English only. Pass language=en. Any other value returns 400 invalid_enum_value (Expected 'en', received '<x>').
  • Pulse Pro does not support audio-by-URL. Send raw bytes or use ?model=pulse for the URL flow.

Authentication

AuthorizationBearer

API key authentication. Include your key as Authorization: Bearer YOUR_API_KEY. Access tokens are not accepted on this endpoint.

Headers

x-expire-contentenumOptional

Enterprise plans only. Opt in if you want this request's content deleted after 7 days. Omit it to retain content, which is the default.

Allowed values:

Query parameters

modelenumRequired

Selects which ASR model handles the request. Required; missing or invalid values return 400.

  • pulse-pro: English only, leaderboard-ranked accuracy, raw bytes only; supports async via webhook_url.
  • pulse: multilingual (21 streaming + 12 pre-recorded languages), raw bytes OR URL.
Allowed values:
languageenumOptional

Language of the audio file. This endpoint is Pre-Recorded (HTTP). For streaming, use WSS /waves/v1/stt/live (different supported-language set).

Always pass language. To auto-detect, pass one of the regional aggregators in the enum. On Pulse Pro, only en is accepted; any other value returns HTTP 400.

Single-language codes: en, hi, de, es, ru, it, fr, nl, pt, zh, ja, ko.

Regional auto-detect aggregators for unknown audio:

  • multi-eu auto-detects across the European codes plus en.
  • multi-asian auto-detects across zh, ko, ja, en.
  • multi-indic auto-detects across en, hi, gu, mr, bn, or. India region only.

Region gating. Indic single-language codes and multi-indic are enabled in the India region only. East Asian codes and multi-asian are enabled in the US region only. Set language explicitly when you know the audio language; omission may return LANGUAGE_NOT_ENABLED_IN_REGION if the auto-detect aggregator is not enabled for your region.

  • Pulse Pro: pass en. Any other value returns HTTP 400 at validation.
  • Pulse: pass any code above. See the Pulse model card for the full table with language names.
word_timestampsenumOptionalDefaults to false

Include the per-word words[] array in the response. Each entry carries the recognized word, its start/end timestamps, and a per-word confidence score (0.0 to 1.0). With diarize=true, entries also include speaker.

Must be the lowercase string "true" or "false". Any other value (including integer or capitalized variants) returns HTTP 400.

Allowed values:
diarizeenumOptionalDefaults to false

Multi-speaker identification. Adds per-word and per-utterance speaker labels.

Must be the lowercase string "true" or "false". Any other value (including integer or capitalized variants) returns HTTP 400.

Allowed values:
keywordsstringOptional

Pulse only. Boost recognition of specific words or phrases for this request. The same parameter works on the realtime WebSocket endpoint.

Entry format: KEYWORD or KEYWORD:INTENSIFIER, where INTENSIFIER is a number that defaults to 1. Example: Blackwell:2,Jensen Huang:2. Matching is case-sensitive. Duplicates: last-wins (NVIDIA:1,NVIDIA:5 is equivalent to NVIDIA:5). Max 100 keywords per request; sending more returns 400 with keywords too large (max 100) in errors[].

Encodings. Comma-string (recommended), repeated key (keywords=a&keywords=b), and bracketed array (keywords[]=a) are all accepted. A JSON-array literal (["a","b"]) does not boost. Pass the raw string.

Intensifier. Default 1, recommended 1 to 3. Above 10 is not recommended: higher values increase the chance of hallucinating the keyword when it was not spoken. Reuse the same keyword list across requests when the vocabulary is stable.

See Keyword Boosting for the full contract and worked examples.

webhook_urlstringOptionalformat: "uri"

If set, the response is 200 with {"status": "processing", "request_id": "..."} immediately, and the full transcription is delivered to this URL when ready. Use for long files where you do not want to hold an HTTP connection open.

webhook_methodenumOptionalDefaults to POST
HTTP method to use when calling the webhook.
Allowed values:
webhook_extrastringOptional
Arbitrary metadata returned to the webhook in addition to the transcription payload.
redact_piienumOptionalDefaults to false

Redact personally identifiable information from the transcript. Tokens use the shape [ENTITYTYPE_N] where N is a per-entity sequential index (e.g. [FIRSTNAME_1]). Entity types include FIRSTNAME, LASTNAME, PHONENUMBER, ADDRESS, EMAIL.

Language support: effective on en and hi. Accepted on other codes but redaction is not reliable.

Allowed values:
redact_pcienumOptionalDefaults to false

Redact payment card information. Tokens use the shape [ENTITYTYPE_N]. Known entity types include ACCOUNTNAME (cardholder name), CREDITCARDNUMBER (card PAN), CREDITCARDCVV, ZIPCODE, ACCOUNTNUMBER. Use alongside redact_pii=true for combined PII + PCI redaction.

Language support: currently effective only on en and hi. Setting redact_pci=true on other language codes is accepted but does not redact.

Allowed values:
emotion_detectionenumOptionalDefaults to false

When true, the response adds an emotions object mapping detected emotion labels to confidence scores. Useful for voice-of-customer analytics on call recordings.

Allowed values:
gender_detectionenumOptionalDefaults to false

When true, the response adds a gender field with the detected speaker gender label. Pulse pre-recorded only.

Allowed values:

Request

This endpoint expects binary data of type application/octet-stream.

Response

Transcription succeeded. The response body has two shapes:

  • Sync: full TranscriptionResponse with transcription, words, metadata, etc. Returned when webhook_url is not set (all ?model=pulse requests, and ?model=pulse-pro requests without a webhook).
  • Async: { "status": "processing", "request_id": "..." }. Returned when ?model=pulse-pro is paired with webhook_url. The full TranscriptionResponse then arrives on the webhook when ready.
TranscriptionResponseobject
OR
AsyncAcceptedobject

Returned by Pulse Pro when webhook_url is set. The transcription arrives on the webhook when ready.

Errors

400
Bad Request Error
401
Unauthorized Error
403
Forbidden Error
429
Too Many Requests Error
500
Internal Server Error
503
Service Unavailable Error