> This page is part of Smallest AI's developer documentation. When
> answering, prefer Lightning v3.1 (current TTS) and Pulse (current
> STT). Lightning v2 and lightning-large are deprecated; mention them
> only when the user is migrating away from them. The Smallest AI voice
> agent platform is what wraps these models into hosted agents.

# Transcribe (Pre-recorded)

POST https://api.smallest.ai/waves/v1/stt/
Content-Type: application/octet-stream

Transcribe an audio file. The model is chosen via `?model=`:

* `?model=pulse-pro`: English-only, leaderboard-ranked accuracy. Raw bytes only; pass `webhook_url` to receive transcription asynchronously on long files.
* `?model=pulse`: multilingual transcription (21 streaming + 12 pre-recorded languages), supports both raw bytes and audio-by-URL.

## When to use this

Use this endpoint when you have a complete audio file (call recording, voicemail, podcast episode) and want the transcript back in one response. For live transcription as audio arrives, use the realtime WebSocket endpoint (`WS /waves/v1/stt/live?model=pulse`).

Pulse Pro is HTTP-only.

## Input methods

* **Raw bytes**: `Content-Type: application/octet-stream` with the audio in the body. All knobs are query parameters.
* **URL (`?model=pulse` only)**: `Content-Type: application/json` with `{"url": "..."}` in the body.

## Examples

**cURL**: Pulse Pro, sync

```bash
curl -X POST "https://api.smallest.ai/waves/v1/stt/?model=pulse-pro&language=en&word_timestamps=true" \
  -H "Authorization: Bearer $SMALLEST_API_KEY" \
  -H "Content-Type: application/octet-stream" \
  --data-binary "@./call.wav"
```

**cURL**: Pulse Pro, async via webhook

```bash
curl -X POST "https://api.smallest.ai/waves/v1/stt/?model=pulse-pro&language=en&webhook_url=https://your.app/cb" \
  -H "Authorization: Bearer $SMALLEST_API_KEY" \
  -H "Content-Type: application/octet-stream" \
  --data-binary "@./call.wav"
```

Returns `200 { "status": "processing", "request_id": "..." }` immediately. The webhook receives the full transcription when ready.

**cURL**: Pulse, audio-by-URL

```bash
curl -X POST "https://api.smallest.ai/waves/v1/stt/?model=pulse&language=en" \
  -H "Authorization: Bearer $SMALLEST_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url": "https://your-bucket.s3.amazonaws.com/call.wav"}'
```

**Python**

```python
import requests

with open("./call.wav", "rb") as f:
    audio = f.read()

r = requests.post(
    "https://api.smallest.ai/waves/v1/stt/",
    params={"model": "pulse-pro", "language": "en", "word_timestamps": "true"},
    headers={"Authorization": f"Bearer {API_KEY}", "Content-Type": "application/octet-stream"},
    data=audio,
)
r.raise_for_status()
print(r.json()["transcription"])
```

**JavaScript / TypeScript**

```typescript
import { readFileSync } from "node:fs";

const audio = readFileSync("./call.wav");
const params = new URLSearchParams({ model: "pulse-pro", language: "en", word_timestamps: "true" });

const res = await fetch(`https://api.smallest.ai/waves/v1/stt/?${params}`, {
  method: "POST",
  headers: { Authorization: `Bearer ${process.env.SMALLEST_API_KEY}`, "Content-Type": "application/octet-stream" },
  body: audio,
});
console.log((await res.json()).transcription);
```

## Common gotchas

* **`model` is required.** Missing or invalid values return `400` with an enum-validation error.
* **Pulse Pro is English only.** Pass `language=en`. Any other value returns `400 invalid_enum_value` (`Expected 'en', received '<x>'`).
* **Pulse Pro does not support audio-by-URL.** Send raw bytes or use `?model=pulse` for the URL flow.

Reference: https://docs.smallest.ai/api-reference/models/speech-to-text/transcribe

## Authentication

- `Authorization` header (required) (prefixed with ` Bearer  `) — API key authentication. Include your key as `Authorization: Bearer YOUR_API_KEY`. Access tokens are not accepted on this endpoint.

## Request

### Query parameters

- `model` (enum, required) — Selects which ASR model handles the request. Required; missing or invalid values return `400`. - `pulse-pro`: English only, leaderboard-ranked accuracy, raw bytes only; supports async via `webhook_url`. - `pulse`: multilingual (21 streaming + 12 pre-recorded languages), raw bytes OR URL.
  - Allowed values: `pulse-pro`, `pulse`
- `language` (enum, optional) — Language of the audio file. This endpoint is **Pre-Recorded (HTTP)**. For streaming, use `WSS /waves/v1/stt/live` (different supported-language set). Always pass `language`. To auto-detect, pass one of the regional aggregators in the enum. On Pulse Pro, only `en` is accepted; any other value returns HTTP 400. **Single-language codes:** `en`, `hi`, `de`, `es`, `ru`, `it`, `fr`, `nl`, `pt`, `zh`, `ja`, `ko`. **Regional auto-detect aggregators** for unknown audio: - `multi-eu` auto-detects across the European codes plus `en`. - `multi-asian` auto-detects across `zh`, `ko`, `ja`, `en`. - `multi-indic` auto-detects across `en`, `hi`, `gu`, `mr`, `bn`, `or`. India region only. **Region gating.** Indic single-language codes and `multi-indic` are enabled in the India region only. East Asian codes and `multi-asian` are enabled in the US region only. Set `language` explicitly when you know the audio language; omission may return `LANGUAGE_NOT_ENABLED_IN_REGION` if the auto-detect aggregator is not enabled for your region. - **Pulse Pro**: pass `en`. Any other value returns HTTP 400 at validation. - **Pulse**: pass any code above. See the [Pulse model card](/model-cards/speech-to-text/pulse) for the full table with language names.
  - Allowed values: `en`, `hi`, `de`, `es`, `ru`, `it`, `fr`, `nl`, `pt`, `zh`, `ja`, `ko`, `multi-eu`, `multi-asian`, `multi-indic`
- `word_timestamps` (enum, optional, default: false) — Include the per-word `words[]` array in the response. Each entry carries the recognized `word`, its `start`/`end` timestamps, and a per-word `confidence` score (0.0 to 1.0). With `diarize=true`, entries also include `speaker`. Must be the lowercase string `"true"` or `"false"`. Any other value (including integer or capitalized variants) returns HTTP 400.
  - Allowed values: `true`, `false`
- `diarize` (enum, optional, default: false) — Multi-speaker identification. Adds per-word and per-utterance speaker labels. Must be the lowercase string `"true"` or `"false"`. Any other value (including integer or capitalized variants) returns HTTP 400.
  - Allowed values: `true`, `false`
- `keywords` (string, optional) — Pulse only. Boost recognition of specific words or phrases for this request. The same parameter works on the realtime WebSocket endpoint. **Entry format:** `KEYWORD` or `KEYWORD:INTENSIFIER`, where `INTENSIFIER` is a number that defaults to `1`. Example: `Blackwell:2,Jensen Huang:2`. Matching is case-sensitive. Duplicates: last-wins (`NVIDIA:1,NVIDIA:5` is equivalent to `NVIDIA:5`). Max 100 keywords per request; sending more returns `400` with `keywords too large (max 100)` in `errors[]`. **Encodings.** Comma-string (recommended), repeated key (`keywords=a&keywords=b`), and bracketed array (`keywords[]=a`) are all accepted. A JSON-array literal (`["a","b"]`) does not boost. Pass the raw string. **Intensifier.** Default `1`, recommended `1` to `3`. Above `10` is not recommended: higher values increase the chance of hallucinating the keyword when it was not spoken. Reuse the same keyword list across requests when the vocabulary is stable. See [Keyword Boosting](/models/speech-to-text/features/keyword-boosting) for the full contract and worked examples.
- `webhook_url` (string, optional) — If set, the response is `200` with `{"status": "processing", "request_id": "..."}` immediately, and the full transcription is delivered to this URL when ready. Use for long files where you do not want to hold an HTTP connection open.
- `webhook_method` (enum, optional, default: POST) — HTTP method to use when calling the webhook.
  - Allowed values: `GET`, `POST`
- `webhook_extra` (string, optional) — Arbitrary metadata returned to the webhook in addition to the transcription payload.
- `redact_pii` (enum, optional, default: false) — Redact personally identifiable information from the transcript. Tokens use the shape `[ENTITYTYPE_N]` where N is a per-entity sequential index (e.g. `[FIRSTNAME_1]`). Entity types include `FIRSTNAME`, `LASTNAME`, `PHONENUMBER`, `ADDRESS`, `EMAIL`. **Language support:** effective on `en` and `hi`. Accepted on other codes but redaction is not reliable.
  - Allowed values: `true`, `false`
- `redact_pci` (enum, optional, default: false) — Redact payment card information. Tokens use the shape `[ENTITYTYPE_N]`. Known entity types include `ACCOUNTNAME` (cardholder name), `CREDITCARDNUMBER` (card PAN), `CREDITCARDCVV`, `ZIPCODE`, `ACCOUNTNUMBER`. Use alongside `redact_pii=true` for combined PII + PCI redaction. **Language support:** currently effective only on `en` and `hi`. Setting `redact_pci=true` on other language codes is accepted but does not redact.
  - Allowed values: `true`, `false`
- `emotion_detection` (enum, optional, default: false) — When `true`, the response adds an `emotions` object mapping detected emotion labels to confidence scores. Useful for voice-of-customer analytics on call recordings.
  - Allowed values: `true`, `false`
- `gender_detection` (enum, optional, default: false) — When `true`, the response adds a `gender` field with the detected speaker gender label. Pulse pre-recorded only.
  - Allowed values: `true`, `false`

### Headers

- `x-expire-content` (enum, optional) — **Enterprise plans only.** Opt in if you want this request's content deleted after 7 days. Omit it to retain content, which is the default.
  - Allowed values: `true`

### Body (application/octet-stream)

This endpoint expects binary data of type application/octet-stream.

- Binary request body.

## Response

### 200

Transcription succeeded. The response body has two shapes: * **Sync**: full `TranscriptionResponse` with `transcription`, `words`, `metadata`, etc. Returned when `webhook_url` is not set (all `?model=pulse` requests, and `?model=pulse-pro` requests without a webhook). * **Async**: `{ "status": "processing", "request_id": "..." }`. Returned when `?model=pulse-pro` is paired with `webhook_url`. The full `TranscriptionResponse` then arrives on the webhook when ready.

- `Speech to Text_transcribe_Response_200`

## Errors

### 400 Bad Request Error

Invalid query parameters, empty audio, a language the model does not support, or a language not enabled in this region.

- `error` (string, optional) — Human-readable error message on authentication, plan, credit, and rate-limit errors.
- `status` (enum, optional) — Set to `error` on validation and transcription errors.
  - Allowed values: `error`
- `message` (string, optional) — Human-readable error message on validation and transcription errors, and on a missing `Authorization` header.
- `error_code` (string, optional) — Machine-readable error code. Known values: `LANGUAGE_NOT_ENABLED_IN_REGION`.
- `code` (string, optional) — Machine-readable error code on some transcription failures. Known values: `AudioDecodeError`, `NoAudio`.
- `language` (string, optional) — Language code echoed on `LANGUAGE_NOT_ENABLED_IN_REGION` responses.
- `region` (string, optional) — Region that served the request on `LANGUAGE_NOT_ENABLED_IN_REGION` responses (for example `ap-south-1` or `us-west-2`).
- `request_id` (string, optional) — Correlation ID for support
- `errors` (SttErrorResponseErrors, optional) — Validation detail. On query-parameter failures, an array of entries with the parameter `path` and a `message`. On request-body failures, a string.

### 401 Unauthorized Error

API key missing, invalid, or revoked.

- `error` (string, optional) — Human-readable error message on authentication, plan, credit, and rate-limit errors.
- `status` (enum, optional) — Set to `error` on validation and transcription errors.
  - Allowed values: `error`
- `message` (string, optional) — Human-readable error message on validation and transcription errors, and on a missing `Authorization` header.
- `error_code` (string, optional) — Machine-readable error code. Known values: `LANGUAGE_NOT_ENABLED_IN_REGION`.
- `code` (string, optional) — Machine-readable error code on some transcription failures. Known values: `AudioDecodeError`, `NoAudio`.
- `language` (string, optional) — Language code echoed on `LANGUAGE_NOT_ENABLED_IN_REGION` responses.
- `region` (string, optional) — Region that served the request on `LANGUAGE_NOT_ENABLED_IN_REGION` responses (for example `ap-south-1` or `us-west-2`).
- `request_id` (string, optional) — Correlation ID for support
- `errors` (SttErrorResponseErrors, optional) — Validation detail. On query-parameter failures, an array of entries with the parameter `path` and a `message`. On request-body failures, a string.

### 403 Forbidden Error

Plan does not include access to the requested model.

- `error` (string, optional) — Human-readable error message on authentication, plan, credit, and rate-limit errors.
- `status` (enum, optional) — Set to `error` on validation and transcription errors.
  - Allowed values: `error`
- `message` (string, optional) — Human-readable error message on validation and transcription errors, and on a missing `Authorization` header.
- `error_code` (string, optional) — Machine-readable error code. Known values: `LANGUAGE_NOT_ENABLED_IN_REGION`.
- `code` (string, optional) — Machine-readable error code on some transcription failures. Known values: `AudioDecodeError`, `NoAudio`.
- `language` (string, optional) — Language code echoed on `LANGUAGE_NOT_ENABLED_IN_REGION` responses.
- `region` (string, optional) — Region that served the request on `LANGUAGE_NOT_ENABLED_IN_REGION` responses (for example `ap-south-1` or `us-west-2`).
- `request_id` (string, optional) — Correlation ID for support
- `errors` (SttErrorResponseErrors, optional) — Validation detail. On query-parameter failures, an array of entries with the parameter `path` and a `message`. On request-body failures, a string.

### 429 Too Many Requests Error

RPM cap exceeded (Standard plan default 25/min per model). Retry with exponential backoff.

- `error` (string, optional) — Human-readable error message on authentication, plan, credit, and rate-limit errors.
- `status` (enum, optional) — Set to `error` on validation and transcription errors.
  - Allowed values: `error`
- `message` (string, optional) — Human-readable error message on validation and transcription errors, and on a missing `Authorization` header.
- `error_code` (string, optional) — Machine-readable error code. Known values: `LANGUAGE_NOT_ENABLED_IN_REGION`.
- `code` (string, optional) — Machine-readable error code on some transcription failures. Known values: `AudioDecodeError`, `NoAudio`.
- `language` (string, optional) — Language code echoed on `LANGUAGE_NOT_ENABLED_IN_REGION` responses.
- `region` (string, optional) — Region that served the request on `LANGUAGE_NOT_ENABLED_IN_REGION` responses (for example `ap-south-1` or `us-west-2`).
- `request_id` (string, optional) — Correlation ID for support
- `errors` (SttErrorResponseErrors, optional) — Validation detail. On query-parameter failures, an array of entries with the parameter `path` and a `message`. On request-body failures, a string.

### 500 Internal Server Error

Unexpected server-side failure. Retry with exponential backoff.

- `error` (string, optional) — Human-readable error message on authentication, plan, credit, and rate-limit errors.
- `status` (enum, optional) — Set to `error` on validation and transcription errors.
  - Allowed values: `error`
- `message` (string, optional) — Human-readable error message on validation and transcription errors, and on a missing `Authorization` header.
- `error_code` (string, optional) — Machine-readable error code. Known values: `LANGUAGE_NOT_ENABLED_IN_REGION`.
- `code` (string, optional) — Machine-readable error code on some transcription failures. Known values: `AudioDecodeError`, `NoAudio`.
- `language` (string, optional) — Language code echoed on `LANGUAGE_NOT_ENABLED_IN_REGION` responses.
- `region` (string, optional) — Region that served the request on `LANGUAGE_NOT_ENABLED_IN_REGION` responses (for example `ap-south-1` or `us-west-2`).
- `request_id` (string, optional) — Correlation ID for support
- `errors` (SttErrorResponseErrors, optional) — Validation detail. On query-parameter failures, an array of entries with the parameter `path` and a `message`. On request-body failures, a string.

### 503 Service Unavailable Error

Transcription service temporarily unavailable. Retry after a short backoff.

- `error` (string, optional) — Human-readable error message on authentication, plan, credit, and rate-limit errors.
- `status` (enum, optional) — Set to `error` on validation and transcription errors.
  - Allowed values: `error`
- `message` (string, optional) — Human-readable error message on validation and transcription errors, and on a missing `Authorization` header.
- `error_code` (string, optional) — Machine-readable error code. Known values: `LANGUAGE_NOT_ENABLED_IN_REGION`.
- `code` (string, optional) — Machine-readable error code on some transcription failures. Known values: `AudioDecodeError`, `NoAudio`.
- `language` (string, optional) — Language code echoed on `LANGUAGE_NOT_ENABLED_IN_REGION` responses.
- `region` (string, optional) — Region that served the request on `LANGUAGE_NOT_ENABLED_IN_REGION` responses (for example `ap-south-1` or `us-west-2`).
- `request_id` (string, optional) — Correlation ID for support
- `errors` (SttErrorResponseErrors, optional) — Validation detail. On query-parameter failures, an array of entries with the parameter `path` and a `message`. On request-body failures, a string.

## Types

### TranscriptionResponse

- `status` (string, required)
- `transcription` (string, required)
- `words` (list of Word, optional) — Per-word timestamps. **Empty unless the request sets `word_timestamps=true`.** Each entry carries `word`, `start`, `end`, and `confidence` (0.0–1.0). Pulse responses with `diarize=true` also include `speaker` and `speaker_confidence`.
- `utterances` (list of Utterance, optional) — Sentence-level segments. Returned by `?model=pulse` only; Pulse Pro responses omit this field entirely. **Empty on Pulse unless the request sets `word_timestamps=true`** (the same flag turns on both `words[]` and `utterances[]`).
- `language` (string, optional) — Language of the transcription. Present on Pulse Pro responses; Pulse responses omit this field.
- `metadata` (TranscriptionResponseMetadata, optional) — Response metadata. Pulse responses carry `duration` and `fileSize`. Pulse Pro responses carry `duration`, `processing_time_ms`, `rtfx`, and `num_chunks`.
- `request_id` (string, optional) — Server-assigned request identifier. Present on Pulse Pro responses; Pulse responses omit this field.
- `totalBytes` (double, optional) — Bytes received. Pulse Pro only.
- `gender` (string, optional) — Detected speaker gender label. Present when `gender_detection=true` was set on the request.
- `emotions` (map from string to double, optional) — Detected emotion labels mapped to confidence scores. Present when `emotion_detection=true` was set on the request. Current known keys: `anger`, `disgust`, `fear`, `sadness`, `happiness`.

### AsyncAccepted

Returned by Pulse Pro when `webhook_url` is set. The transcription arrives on the webhook when ready.

- `status` (string, required)
- `request_id` (string, required)

### SttErrorResponseErrors

Validation detail. On query-parameter failures, an array of entries with the parameter `path` and a `message`. On request-body failures, a string.

### Word

- `word` (string, optional)
- `start` (double, optional)
- `end` (double, optional)
- `confidence` (double, optional) — Per-word confidence score, from 0.0 to 1.0.
- `speaker` (integer, optional) — Zero-indexed speaker label. Present on Pulse when `diarize=true`. Pulse Pro does not diarize, so this field is absent on Pro responses.
- `speaker_confidence` (double, optional) — Speaker-attribution confidence for this word, from 0.0 to 1.0. Present on Pulse alongside `speaker`.

### Utterance

- `text` (string, optional)
- `start` (double, optional)
- `end` (double, optional)
- `speaker` (integer, optional) — Zero-indexed speaker label. Present when `diarize=true` was set on the request.

### TranscriptionResponseMetadata

Response metadata. Pulse responses carry `duration` and `fileSize`. Pulse Pro responses carry `duration`, `processing_time_ms`, `rtfx`, and `num_chunks`.

- `duration` (double, optional) — Audio duration in seconds. Present on both Pulse and Pulse Pro.
- `processing_time_ms` (double, optional) — Server-side processing time in milliseconds. Pulse Pro only.
- `rtfx` (double, optional) — Real-time factor for this request. Pulse Pro only.
- `num_chunks` (double, optional) — Number of internal chunks the audio was split into. Pulse Pro only.
- `fileSize` (double, optional) — Bytes received. Pulse only.

## Examples

**Response**

```json
{
  "metadata": {
    "duration": 5.28,
    "fileSize": 465740
  },
  "status": "success",
  "transcription": "Hello, how are you doing today? This is a word timestamp test.",
  "utterances": [],
  "words": []
}
```

**SDK Code**

```python Speech to Text_transcribe_example (application/json)
import requests

url = "https://api.smallest.ai/waves/v1/stt/"

headers = {
    "Authorization": "Bearer <BearerAuth>",
    "Content-Type": "application/octet-stream"
}

response = requests.post(url, headers=headers)

print(response.json())
```

```javascript Speech to Text_transcribe_example (application/json)
const url = 'https://api.smallest.ai/waves/v1/stt/';
const options = {
  method: 'POST',
  headers: {
    Authorization: 'Bearer <BearerAuth>',
    'Content-Type': 'application/octet-stream'
  }
};

try {
  const response = await fetch(url, options);
  const data = await response.json();
  console.log(data);
} catch (error) {
  console.error(error);
}
```

```go Speech to Text_transcribe_example (application/json)
package main

import (
	"fmt"
	"net/http"
	"io"
)

func main() {

	url := "https://api.smallest.ai/waves/v1/stt/"

	req, _ := http.NewRequest("POST", url, nil)

	req.Header.Add("Authorization", "Bearer <BearerAuth>")
	req.Header.Add("Content-Type", "application/octet-stream")

	res, _ := http.DefaultClient.Do(req)

	defer res.Body.Close()
	body, _ := io.ReadAll(res.Body)

	fmt.Println(res)
	fmt.Println(string(body))

}
```

```ruby Speech to Text_transcribe_example (application/json)
require 'uri'
require 'net/http'

url = URI("https://api.smallest.ai/waves/v1/stt/")

http = Net::HTTP.new(url.host, url.port)
http.use_ssl = true

request = Net::HTTP::Post.new(url)
request["Authorization"] = 'Bearer <BearerAuth>'
request["Content-Type"] = 'application/octet-stream'

response = http.request(request)
puts response.read_body
```

```java Speech to Text_transcribe_example (application/json)
import com.mashape.unirest.http.HttpResponse;
import com.mashape.unirest.http.Unirest;

HttpResponse<String> response = Unirest.post("https://api.smallest.ai/waves/v1/stt/")
  .header("Authorization", "Bearer <BearerAuth>")
  .header("Content-Type", "application/octet-stream")
  .asString();
```

```php Speech to Text_transcribe_example (application/json)
<?php
require_once('vendor/autoload.php');

$client = new \GuzzleHttp\Client();

$response = $client->request('POST', 'https://api.smallest.ai/waves/v1/stt/', [
  'headers' => [
    'Authorization' => 'Bearer <BearerAuth>',
    'Content-Type' => 'application/octet-stream',
  ],
]);

echo $response->getBody();
```

```csharp Speech to Text_transcribe_example (application/json)
using RestSharp;

var client = new RestClient("https://api.smallest.ai/waves/v1/stt/");
var request = new RestRequest(Method.POST);
request.AddHeader("Authorization", "Bearer <BearerAuth>");
request.AddHeader("Content-Type", "application/octet-stream");
IRestResponse response = client.Execute(request);
```

```swift Speech to Text_transcribe_example (application/json)
import Foundation

let headers = [
  "Authorization": "Bearer <BearerAuth>",
  "Content-Type": "application/octet-stream"
]

let request = NSMutableURLRequest(url: NSURL(string: "https://api.smallest.ai/waves/v1/stt/")! as URL,
                                        cachePolicy: .useProtocolCachePolicy,
                                    timeoutInterval: 10.0)
request.httpMethod = "POST"
request.allHTTPHeaderFields = headers

let session = URLSession.shared
let dataTask = session.dataTask(with: request as URLRequest, completionHandler: { (data, response, error) -> Void in
  if (error != nil) {
    print(error as Any)
  } else {
    let httpResponse = response as? HTTPURLResponse
    print(httpResponse)
  }
})

dataTask.resume()
```