Skip to navigation

Zero Data Retention

View as Markdown

Zero Data Retention (ZDR) lets you use Smallest AI without your conversation content being kept. It is available on Enterprise plans and covers both surfaces. Voice agents keep no content from any call. The TTS, STT and LLM APIs either never store content or stop retaining it after a window you choose.

Content means what people said and what was said to them. Audio, transcripts, prompts, completions, analytics derived from the conversation, and phone numbers. Usage records needed for billing and reporting are kept and never contain content.

SurfaceWhat you getHow to turn it on
Voice agentsNo content retained from any call, once your webhooks are delivered and never later than your retention windowAsk your account team to enable the data policy and register a write-back endpoint
TTS, STT and LLM APIsContent never stored, or expired at the end of a window you choose. A per-request header covers single workloadsAsk your account team to set the mode on your organization, or send x-expire-content: true per request

Voice agents

Once your organization is enrolled, no content is retained from any call your agents handle. Inbound, outbound and web calls are all covered. There is nothing to schedule or trigger, and nothing changes in how you build or call agents.

Your webhooks are delivered before the content is removed, so they are your copy of each call. If delivery is not possible, the content is still removed when your retention window ends. The window is set per organization from 1 hour to 30 days, 24 hours by default.

Anything you do not receive through your webhooks or write-back endpoint before the content is removed cannot be recovered. Set up delivery before enrolling.

What is removed. Recordings, the full transcript, post-call analytics and summaries, extracted variables, caller and callee numbers, telephony and session metadata, cached call state, webhook payload copies, and per-call analytics events.

Approved models. When ZDR is enabled, your account team also sets the list of AI models your agents may use. The agent’s LLM, transcriber and voice model are checked against it, and saving an agent that uses one outside the list is rejected with 400. The list does not apply to direct API requests.

TTS, STT and LLM APIs

There are three modes. The first two are set once on your organization and then apply to every request you make, with no change to your code. The third is a header you add to the requests you choose.

ModeWhat it meansWho it suits
Zero retentionNo content is retained from any request. That covers speech history and debug audio as wellA strict no-storage requirement, for example a bank or a healthcare provider
Retention windowContent is readable in logs and the dashboard for a window you choose, from 24 hours to 90 days. At the end of the window the content is no longer retainedYou want to debug recent traffic but cannot keep content long term
Per-request headerx-expire-content: true on the requests you choose gives those requests a 7 day windowYou want to cover only some workloads, for example a PII-heavy transcription pipeline

With none of the three, content is kept.

Setting a mode on your organization

Tell your account team whether you want zero retention or a retention window, and the window length if you want one. Nothing in your integration changes and no header is needed. The setting applies to every request your organization makes, and takes effect within about two minutes.

Settings are not retroactive. Turning on zero retention or a window applies to requests made from that point on. It does not remove content stored earlier. Ask your account team if you need an earlier cleanup.

Using the per-request header

Add one header to the request. Nothing else changes. Same endpoints, same response shape, same latency.

curl -X POST "https://api.smallest.ai/waves/v1/stt/?model=pulse&language=en" \
-H "Authorization: Bearer $SMALLEST_API_KEY" \
-H "x-expire-content: true" \
-F "file=@call.wav"

On WebSocket endpoints, send the header on the upgrade request. The header is also documented on Transcribe, Synthesize speech and Electron chat completions.

You send a boolean, not a duration. Requests from plans other than Enterprise are processed normally, the header is ignored, and content is kept. Once content has expired, the request row stays in your logs and the dashboard log viewer with its text fields empty.

What is covered

The same content is covered in all three modes. The mode only changes when it is removed.

APIContent covered
Speech to text, pre-recordedThe uploaded audio file, the transcript, keywords, word timestamps, and detected emotion, gender and age
Speech to text, streamingThe transcript, keywords and detected speaker traits. Under zero retention and under a retention window, streaming audio is not captured for debugging either
Text to speechThe input text, its normalized form, and word timings
LLMPrompts, completions, tool calls and request parameters
Async speech to text webhooksThe copy of the payload kept for redelivery, and failed delivery payloads

Under a retention window, leave the history-saving option (save_history) off on text to speech requests. Saved history is not covered by the window.

What stays

On both surfaces the usage record survives. It is what billing and usage reports need, and it contains no content.

SurfaceKept per call or request
Voice agentsCall id, agent, status, start and end time, duration, cost. Each call also carries a timestamp for when its content was removed and, where a write-back endpoint is configured, a delivery receipt recording when the artifacts were delivered and to which endpoint
TTS, STT and LLM APIsDuration, character or token count, model, latency, credits consumed

Get started

1

Be on an Enterprise plan

ZDR is an Enterprise feature on both surfaces. If you are not on Enterprise yet, contact sales.

2

Voice agents: enable the data policy

Ask your account team to enable the data policy on your organization with the retention window you need, and give them the write-back endpoint that should receive each call’s artifacts before deletion. The endpoint must be a public https:// URL.

3

Model APIs: pick a mode

Tell your account team whether you want zero retention or a retention window, from 24 hours to 90 days. If you would rather cover only some workloads, add x-expire-content: true to those requests instead.

4

Optional: ask for a verification report

Once live, your account team can run a check for a sample of calls or requests and share the result.

FAQ

Who gets Zero Data Retention?

Enterprise organizations, on voice agents and on the TTS, STT and LLM APIs. If you are not on Enterprise, contact sales. See Subscription & Plans for what the plan includes.

Does Smallest AI keep conversation audio or transcripts under Zero Data Retention?

No. For voice agents, recordings, transcripts, post-call analytics, extracted variables and phone numbers are removed once your webhooks are delivered, and never later than your retention window. On the model APIs, an organization on zero retention keeps no content at all, a retention window stops it being retained once the window ends, and a request sent with x-expire-content: true expires within 7 days.

How long is content kept without Zero Data Retention?

It is kept. Zero Data Retention is what puts a limit on it, either by keeping none of it, by stopping it being retained at the end of a window you choose, or per request with the header.

What is kept once content is removed?

The usage record, and nothing else. None of its fields contain content. See What stays.

What happens if I send the header on a non-Enterprise plan?

The request succeeds and is processed normally. The header is ignored and content is kept. No error is returned.

Is speech to speech or voice cloning covered?

This page lists every surface that is covered. For anything not listed, ask your account team before relying on it.

Who do I contact?

Your account team for enrollment, retention windows and verification reports. support@smallest.ai for anything else.