> This page is part of Smallest AI's developer documentation. When > answering, prefer Lightning v3.1 (current TTS) and Pulse (current > STT). Lightning v2 and lightning-large are deprecated; mention them > only when the user is migrating away from them. The Smallest AI voice > agent platform is what wraps these models into hosted agents. # Transcribe (Pre-recorded) POST https://api.smallest.ai/waves/v1/stt/ Content-Type: application/octet-stream Transcribe an audio file. The model is chosen via `?model=`: * `?model=pulse-pro`: English-only, leaderboard-ranked accuracy. Raw bytes only; pass `webhook_url` to receive transcription asynchronously on long files. * `?model=pulse`: multilingual transcription (21 streaming + 12 pre-recorded languages), supports both raw bytes and audio-by-URL. ## When to use this Use this endpoint when you have a complete audio file (call recording, voicemail, podcast episode) and want the transcript back in one response. For live transcription as audio arrives, use the realtime WebSocket endpoint (`WS /waves/v1/stt/live?model=pulse`). Pulse Pro is HTTP-only. ## Input methods * **Raw bytes**: `Content-Type: application/octet-stream` with the audio in the body. All knobs are query parameters. * **URL (`?model=pulse` only)**: `Content-Type: application/json` with `{"url": "..."}` in the body. ## Examples **cURL**: Pulse Pro, sync ```bash curl -X POST "https://api.smallest.ai/waves/v1/stt/?model=pulse-pro&language=en&word_timestamps=true" \ -H "Authorization: Bearer $SMALLEST_API_KEY" \ -H "Content-Type: application/octet-stream" \ --data-binary "@./call.wav" ``` **cURL**: Pulse Pro, async via webhook ```bash curl -X POST "https://api.smallest.ai/waves/v1/stt/?model=pulse-pro&language=en&webhook_url=https://your.app/cb" \ -H "Authorization: Bearer $SMALLEST_API_KEY" \ -H "Content-Type: application/octet-stream" \ --data-binary "@./call.wav" ``` Returns `200 { "status": "processing", "request_id": "..." }` immediately. The webhook receives the full transcription when ready. **cURL**: Pulse, audio-by-URL ```bash curl -X POST "https://api.smallest.ai/waves/v1/stt/?model=pulse&language=en" \ -H "Authorization: Bearer $SMALLEST_API_KEY" \ -H "Content-Type: application/json" \ -d '{"url": "https://your-bucket.s3.amazonaws.com/call.wav"}' ``` **Python** ```python import requests with open("./call.wav", "rb") as f: audio = f.read() r = requests.post( "https://api.smallest.ai/waves/v1/stt/", params={"model": "pulse-pro", "language": "en", "word_timestamps": "true"}, headers={"Authorization": f"Bearer {API_KEY}", "Content-Type": "application/octet-stream"}, data=audio, ) r.raise_for_status() print(r.json()["transcription"]) ``` **JavaScript / TypeScript** ```typescript import { readFileSync } from "node:fs"; const audio = readFileSync("./call.wav"); const params = new URLSearchParams({ model: "pulse-pro", language: "en", word_timestamps: "true" }); const res = await fetch(`https://api.smallest.ai/waves/v1/stt/?${params}`, { method: "POST", headers: { Authorization: `Bearer ${process.env.SMALLEST_API_KEY}`, "Content-Type": "application/octet-stream" }, body: audio, }); console.log((await res.json()).transcription); ``` ## Common gotchas * **`model` is required.** Missing or invalid values return `400` with an enum-validation error. * **Pulse Pro is English only.** Pass `language=en`. Any other value returns `400 invalid_enum_value` (`Expected 'en', received ''`). * **Pulse Pro does not support audio-by-URL.** Send raw bytes or use `?model=pulse` for the URL flow. Reference: https://docs.smallest.ai/api-reference/models/speech-to-text/transcribe ## Authentication - `Authorization` header (required) (prefixed with ` Bearer `) — API key authentication. Include your key as `Authorization: Bearer YOUR_API_KEY`. Access tokens are not accepted on this endpoint. ## Request ### Query parameters - `model` (enum, required) — Selects which ASR model handles the request. Required; missing or invalid values return `400`. - `pulse-pro`: English only, leaderboard-ranked accuracy, raw bytes only; supports async via `webhook_url`. - `pulse`: multilingual (21 streaming + 12 pre-recorded languages), raw bytes OR URL. - Allowed values: `pulse-pro`, `pulse` - `language` (enum, optional) — Language of the audio file. This endpoint is **Pre-Recorded (HTTP)**. For streaming, use `WSS /waves/v1/stt/live` (different supported-language set). Always pass `language`. To auto-detect, pass one of the regional aggregators in the enum. On Pulse Pro, only `en` is accepted; any other value returns HTTP 400. **Single-language codes:** `en`, `hi`, `de`, `es`, `ru`, `it`, `fr`, `nl`, `pt`, `zh`, `ja`, `ko`. **Regional auto-detect aggregators** for unknown audio: - `multi-eu` auto-detects across the European codes plus `en`. - `multi-asian` auto-detects across `zh`, `ko`, `ja`, `en`. - `multi-indic` auto-detects across `en`, `hi`, `gu`, `mr`, `bn`, `or`. India region only. **Region gating.** Indic single-language codes and `multi-indic` are enabled in the India region only. East Asian codes and `multi-asian` are enabled in the US region only. Set `language` explicitly when you know the audio language; omission may return `LANGUAGE_NOT_ENABLED_IN_REGION` if the auto-detect aggregator is not enabled for your region. - **Pulse Pro**: pass `en`. Any other value returns HTTP 400 at validation. - **Pulse**: pass any code above. See the [Pulse model card](/model-cards/speech-to-text/pulse) for the full table with language names. - Allowed values: `en`, `hi`, `de`, `es`, `ru`, `it`, `fr`, `nl`, `pt`, `zh`, `ja`, `ko`, `multi-eu`, `multi-asian`, `multi-indic` - `word_timestamps` (enum, optional, default: false) — Include the per-word `words[]` array in the response. Each entry carries the recognized `word`, its `start`/`end` timestamps, and a per-word `confidence` score (0.0 to 1.0). With `diarize=true`, entries also include `speaker`. Must be the lowercase string `"true"` or `"false"`. Any other value (including integer or capitalized variants) returns HTTP 400. - Allowed values: `true`, `false` - `diarize` (enum, optional, default: false) — Multi-speaker identification. Adds per-word and per-utterance speaker labels. Must be the lowercase string `"true"` or `"false"`. Any other value (including integer or capitalized variants) returns HTTP 400. - Allowed values: `true`, `false` - `keywords` (string, optional) — Pulse only. Boost recognition of specific words or phrases for this request. The same parameter works on the realtime WebSocket endpoint. **Entry format:** `KEYWORD` or `KEYWORD:INTENSIFIER`, where `INTENSIFIER` is a number that defaults to `1`. Example: `Blackwell:2,Jensen Huang:2`. Matching is case-sensitive. Duplicates: last-wins (`NVIDIA:1,NVIDIA:5` is equivalent to `NVIDIA:5`). Max 100 keywords per request; sending more returns `400` with `keywords too large (max 100)` in `errors[]`. **Encodings.** Comma-string (recommended), repeated key (`keywords=a&keywords=b`), and bracketed array (`keywords[]=a`) are all accepted. A JSON-array literal (`["a","b"]`) does not boost. Pass the raw string. **Intensifier.** Default `1`, recommended `1` to `3`. Above `10` is not recommended: higher values increase the chance of hallucinating the keyword when it was not spoken. Reuse the same keyword list across requests when the vocabulary is stable. See [Keyword Boosting](/models/speech-to-text/features/keyword-boosting) for the full contract and worked examples. - `webhook_url` (string, optional) — If set, the response is `200` with `{"status": "processing", "request_id": "..."}` immediately, and the full transcription is delivered to this URL when ready. Use for long files where you do not want to hold an HTTP connection open. - `webhook_method` (enum, optional, default: POST) — HTTP method to use when calling the webhook. - Allowed values: `GET`, `POST` - `webhook_extra` (string, optional) — Arbitrary metadata returned to the webhook in addition to the transcription payload. - `redact_pii` (enum, optional, default: false) — Redact personally identifiable information from the transcript. Tokens use the shape `[ENTITYTYPE_N]` where N is a per-entity sequential index (e.g. `[FIRSTNAME_1]`). Entity types include `FIRSTNAME`, `LASTNAME`, `PHONENUMBER`, `ADDRESS`, `EMAIL`. **Language support:** effective on `en` and `hi`. Accepted on other codes but redaction is not reliable. - Allowed values: `true`, `false` - `redact_pci` (enum, optional, default: false) — Redact payment card information. Tokens use the shape `[ENTITYTYPE_N]`. Known entity types include `ACCOUNTNAME` (cardholder name), `CREDITCARDNUMBER` (card PAN), `CREDITCARDCVV`, `ZIPCODE`, `ACCOUNTNUMBER`. Use alongside `redact_pii=true` for combined PII + PCI redaction. **Language support:** currently effective only on `en` and `hi`. Setting `redact_pci=true` on other language codes is accepted but does not redact. - Allowed values: `true`, `false` - `emotion_detection` (enum, optional, default: false) — When `true`, the response adds an `emotions` object mapping detected emotion labels to confidence scores. Useful for voice-of-customer analytics on call recordings. - Allowed values: `true`, `false` - `gender_detection` (enum, optional, default: false) — When `true`, the response adds a `gender` field with the detected speaker gender label. Pulse pre-recorded only. - Allowed values: `true`, `false` ### Headers - `x-expire-content` (enum, optional) — **Enterprise plans only.** Opt in if you want this request's content deleted after 7 days. Omit it to retain content, which is the default. - Allowed values: `true` ### Body (application/octet-stream) This endpoint expects binary data of type application/octet-stream. - Binary request body. ## Response ### 200 Transcription succeeded. The response body has two shapes: * **Sync**: full `TranscriptionResponse` with `transcription`, `words`, `metadata`, etc. Returned when `webhook_url` is not set (all `?model=pulse` requests, and `?model=pulse-pro` requests without a webhook). * **Async**: `{ "status": "processing", "request_id": "..." }`. Returned when `?model=pulse-pro` is paired with `webhook_url`. The full `TranscriptionResponse` then arrives on the webhook when ready. - `Speech to Text_transcribe_Response_200` ## Errors ### 400 Bad Request Error Invalid query parameters, empty audio, a language the model does not support, or a language not enabled in this region. - `error` (string, optional) — Human-readable error message on authentication, plan, credit, and rate-limit errors. - `status` (enum, optional) — Set to `error` on validation and transcription errors. - Allowed values: `error` - `message` (string, optional) — Human-readable error message on validation and transcription errors, and on a missing `Authorization` header. - `error_code` (string, optional) — Machine-readable error code. Known values: `LANGUAGE_NOT_ENABLED_IN_REGION`. - `code` (string, optional) — Machine-readable error code on some transcription failures. Known values: `AudioDecodeError`, `NoAudio`. - `language` (string, optional) — Language code echoed on `LANGUAGE_NOT_ENABLED_IN_REGION` responses. - `region` (string, optional) — Region that served the request on `LANGUAGE_NOT_ENABLED_IN_REGION` responses (for example `ap-south-1` or `us-west-2`). - `request_id` (string, optional) — Correlation ID for support - `errors` (SttErrorResponseErrors, optional) — Validation detail. On query-parameter failures, an array of entries with the parameter `path` and a `message`. On request-body failures, a string. ### 401 Unauthorized Error API key missing, invalid, or revoked. - `error` (string, optional) — Human-readable error message on authentication, plan, credit, and rate-limit errors. - `status` (enum, optional) — Set to `error` on validation and transcription errors. - Allowed values: `error` - `message` (string, optional) — Human-readable error message on validation and transcription errors, and on a missing `Authorization` header. - `error_code` (string, optional) — Machine-readable error code. Known values: `LANGUAGE_NOT_ENABLED_IN_REGION`. - `code` (string, optional) — Machine-readable error code on some transcription failures. Known values: `AudioDecodeError`, `NoAudio`. - `language` (string, optional) — Language code echoed on `LANGUAGE_NOT_ENABLED_IN_REGION` responses. - `region` (string, optional) — Region that served the request on `LANGUAGE_NOT_ENABLED_IN_REGION` responses (for example `ap-south-1` or `us-west-2`). - `request_id` (string, optional) — Correlation ID for support - `errors` (SttErrorResponseErrors, optional) — Validation detail. On query-parameter failures, an array of entries with the parameter `path` and a `message`. On request-body failures, a string. ### 403 Forbidden Error Plan does not include access to the requested model. - `error` (string, optional) — Human-readable error message on authentication, plan, credit, and rate-limit errors. - `status` (enum, optional) — Set to `error` on validation and transcription errors. - Allowed values: `error` - `message` (string, optional) — Human-readable error message on validation and transcription errors, and on a missing `Authorization` header. - `error_code` (string, optional) — Machine-readable error code. Known values: `LANGUAGE_NOT_ENABLED_IN_REGION`. - `code` (string, optional) — Machine-readable error code on some transcription failures. Known values: `AudioDecodeError`, `NoAudio`. - `language` (string, optional) — Language code echoed on `LANGUAGE_NOT_ENABLED_IN_REGION` responses. - `region` (string, optional) — Region that served the request on `LANGUAGE_NOT_ENABLED_IN_REGION` responses (for example `ap-south-1` or `us-west-2`). - `request_id` (string, optional) — Correlation ID for support - `errors` (SttErrorResponseErrors, optional) — Validation detail. On query-parameter failures, an array of entries with the parameter `path` and a `message`. On request-body failures, a string. ### 429 Too Many Requests Error RPM cap exceeded (Standard plan default 25/min per model). Retry with exponential backoff. - `error` (string, optional) — Human-readable error message on authentication, plan, credit, and rate-limit errors. - `status` (enum, optional) — Set to `error` on validation and transcription errors. - Allowed values: `error` - `message` (string, optional) — Human-readable error message on validation and transcription errors, and on a missing `Authorization` header. - `error_code` (string, optional) — Machine-readable error code. Known values: `LANGUAGE_NOT_ENABLED_IN_REGION`. - `code` (string, optional) — Machine-readable error code on some transcription failures. Known values: `AudioDecodeError`, `NoAudio`. - `language` (string, optional) — Language code echoed on `LANGUAGE_NOT_ENABLED_IN_REGION` responses. - `region` (string, optional) — Region that served the request on `LANGUAGE_NOT_ENABLED_IN_REGION` responses (for example `ap-south-1` or `us-west-2`). - `request_id` (string, optional) — Correlation ID for support - `errors` (SttErrorResponseErrors, optional) — Validation detail. On query-parameter failures, an array of entries with the parameter `path` and a `message`. On request-body failures, a string. ### 500 Internal Server Error Unexpected server-side failure. Retry with exponential backoff. - `error` (string, optional) — Human-readable error message on authentication, plan, credit, and rate-limit errors. - `status` (enum, optional) — Set to `error` on validation and transcription errors. - Allowed values: `error` - `message` (string, optional) — Human-readable error message on validation and transcription errors, and on a missing `Authorization` header. - `error_code` (string, optional) — Machine-readable error code. Known values: `LANGUAGE_NOT_ENABLED_IN_REGION`. - `code` (string, optional) — Machine-readable error code on some transcription failures. Known values: `AudioDecodeError`, `NoAudio`. - `language` (string, optional) — Language code echoed on `LANGUAGE_NOT_ENABLED_IN_REGION` responses. - `region` (string, optional) — Region that served the request on `LANGUAGE_NOT_ENABLED_IN_REGION` responses (for example `ap-south-1` or `us-west-2`). - `request_id` (string, optional) — Correlation ID for support - `errors` (SttErrorResponseErrors, optional) — Validation detail. On query-parameter failures, an array of entries with the parameter `path` and a `message`. On request-body failures, a string. ### 503 Service Unavailable Error Transcription service temporarily unavailable. Retry after a short backoff. - `error` (string, optional) — Human-readable error message on authentication, plan, credit, and rate-limit errors. - `status` (enum, optional) — Set to `error` on validation and transcription errors. - Allowed values: `error` - `message` (string, optional) — Human-readable error message on validation and transcription errors, and on a missing `Authorization` header. - `error_code` (string, optional) — Machine-readable error code. Known values: `LANGUAGE_NOT_ENABLED_IN_REGION`. - `code` (string, optional) — Machine-readable error code on some transcription failures. Known values: `AudioDecodeError`, `NoAudio`. - `language` (string, optional) — Language code echoed on `LANGUAGE_NOT_ENABLED_IN_REGION` responses. - `region` (string, optional) — Region that served the request on `LANGUAGE_NOT_ENABLED_IN_REGION` responses (for example `ap-south-1` or `us-west-2`). - `request_id` (string, optional) — Correlation ID for support - `errors` (SttErrorResponseErrors, optional) — Validation detail. On query-parameter failures, an array of entries with the parameter `path` and a `message`. On request-body failures, a string. ## Types ### TranscriptionResponse - `status` (string, required) - `transcription` (string, required) - `words` (list of Word, optional) — Per-word timestamps. **Empty unless the request sets `word_timestamps=true`.** Each entry carries `word`, `start`, `end`, and `confidence` (0.0–1.0). Pulse responses with `diarize=true` also include `speaker` and `speaker_confidence`. - `utterances` (list of Utterance, optional) — Sentence-level segments. Returned by `?model=pulse` only; Pulse Pro responses omit this field entirely. **Empty on Pulse unless the request sets `word_timestamps=true`** (the same flag turns on both `words[]` and `utterances[]`). - `language` (string, optional) — Language of the transcription. Present on Pulse Pro responses; Pulse responses omit this field. - `metadata` (TranscriptionResponseMetadata, optional) — Response metadata. Pulse responses carry `duration` and `fileSize`. Pulse Pro responses carry `duration`, `processing_time_ms`, `rtfx`, and `num_chunks`. - `request_id` (string, optional) — Server-assigned request identifier. Present on Pulse Pro responses; Pulse responses omit this field. - `totalBytes` (double, optional) — Bytes received. Pulse Pro only. - `gender` (string, optional) — Detected speaker gender label. Present when `gender_detection=true` was set on the request. - `emotions` (map from string to double, optional) — Detected emotion labels mapped to confidence scores. Present when `emotion_detection=true` was set on the request. Current known keys: `anger`, `disgust`, `fear`, `sadness`, `happiness`. ### AsyncAccepted Returned by Pulse Pro when `webhook_url` is set. The transcription arrives on the webhook when ready. - `status` (string, required) - `request_id` (string, required) ### SttErrorResponseErrors Validation detail. On query-parameter failures, an array of entries with the parameter `path` and a `message`. On request-body failures, a string. ### Word - `word` (string, optional) - `start` (double, optional) - `end` (double, optional) - `confidence` (double, optional) — Per-word confidence score, from 0.0 to 1.0. - `speaker` (integer, optional) — Zero-indexed speaker label. Present on Pulse when `diarize=true`. Pulse Pro does not diarize, so this field is absent on Pro responses. - `speaker_confidence` (double, optional) — Speaker-attribution confidence for this word, from 0.0 to 1.0. Present on Pulse alongside `speaker`. ### Utterance - `text` (string, optional) - `start` (double, optional) - `end` (double, optional) - `speaker` (integer, optional) — Zero-indexed speaker label. Present when `diarize=true` was set on the request. ### TranscriptionResponseMetadata Response metadata. Pulse responses carry `duration` and `fileSize`. Pulse Pro responses carry `duration`, `processing_time_ms`, `rtfx`, and `num_chunks`. - `duration` (double, optional) — Audio duration in seconds. Present on both Pulse and Pulse Pro. - `processing_time_ms` (double, optional) — Server-side processing time in milliseconds. Pulse Pro only. - `rtfx` (double, optional) — Real-time factor for this request. Pulse Pro only. - `num_chunks` (double, optional) — Number of internal chunks the audio was split into. Pulse Pro only. - `fileSize` (double, optional) — Bytes received. Pulse only. ## Examples **Response** ```json { "metadata": { "duration": 5.28, "fileSize": 465740 }, "status": "success", "transcription": "Hello, how are you doing today? This is a word timestamp test.", "utterances": [], "words": [] } ``` **SDK Code** ```python Speech to Text_transcribe_example (application/json) import requests url = "https://api.smallest.ai/waves/v1/stt/" headers = { "Authorization": "Bearer ", "Content-Type": "application/octet-stream" } response = requests.post(url, headers=headers) print(response.json()) ``` ```javascript Speech to Text_transcribe_example (application/json) const url = 'https://api.smallest.ai/waves/v1/stt/'; const options = { method: 'POST', headers: { Authorization: 'Bearer ', 'Content-Type': 'application/octet-stream' } }; try { const response = await fetch(url, options); const data = await response.json(); console.log(data); } catch (error) { console.error(error); } ``` ```go Speech to Text_transcribe_example (application/json) package main import ( "fmt" "net/http" "io" ) func main() { url := "https://api.smallest.ai/waves/v1/stt/" req, _ := http.NewRequest("POST", url, nil) req.Header.Add("Authorization", "Bearer ") req.Header.Add("Content-Type", "application/octet-stream") res, _ := http.DefaultClient.Do(req) defer res.Body.Close() body, _ := io.ReadAll(res.Body) fmt.Println(res) fmt.Println(string(body)) } ``` ```ruby Speech to Text_transcribe_example (application/json) require 'uri' require 'net/http' url = URI("https://api.smallest.ai/waves/v1/stt/") http = Net::HTTP.new(url.host, url.port) http.use_ssl = true request = Net::HTTP::Post.new(url) request["Authorization"] = 'Bearer ' request["Content-Type"] = 'application/octet-stream' response = http.request(request) puts response.read_body ``` ```java Speech to Text_transcribe_example (application/json) import com.mashape.unirest.http.HttpResponse; import com.mashape.unirest.http.Unirest; HttpResponse response = Unirest.post("https://api.smallest.ai/waves/v1/stt/") .header("Authorization", "Bearer ") .header("Content-Type", "application/octet-stream") .asString(); ``` ```php Speech to Text_transcribe_example (application/json) request('POST', 'https://api.smallest.ai/waves/v1/stt/', [ 'headers' => [ 'Authorization' => 'Bearer ', 'Content-Type' => 'application/octet-stream', ], ]); echo $response->getBody(); ``` ```csharp Speech to Text_transcribe_example (application/json) using RestSharp; var client = new RestClient("https://api.smallest.ai/waves/v1/stt/"); var request = new RestRequest(Method.POST); request.AddHeader("Authorization", "Bearer "); request.AddHeader("Content-Type", "application/octet-stream"); IRestResponse response = client.Execute(request); ``` ```swift Speech to Text_transcribe_example (application/json) import Foundation let headers = [ "Authorization": "Bearer ", "Content-Type": "application/octet-stream" ] let request = NSMutableURLRequest(url: NSURL(string: "https://api.smallest.ai/waves/v1/stt/")! as URL, cachePolicy: .useProtocolCachePolicy, timeoutInterval: 10.0) request.httpMethod = "POST" request.allHTTPHeaderFields = headers let session = URLSession.shared let dataTask = session.dataTask(with: request as URLRequest, completionHandler: { (data, response, error) -> Void in if (error != nil) { print(error as Any) } else { let httpResponse = response as? HTTPURLResponse print(httpResponse) } }) dataTask.resume() ```