Agora

View as Markdown

This guide walks you through configuring Smallest AI as the speech-to-text and text-to-speech provider in Agora Conversational AI Engine. Conversational AI Engine is Agora’s managed pipeline (RTC transport → ASR → LLM → TTS) for building real-time voice agents - vendors are selected declaratively in the agent’s join config, so switching to Smallest AI is a configuration change, not a code change.

Smallest AI is supported as both the ASR and TTS vendor:


Prerequisites

  • An Agora App ID and a Conversational AI Engine-enabled project
  • A Smallest AI API key - get one from the Smallest AI dashboard
  • An LLM provider key (e.g. OpenAI) for the llm leg of the pipeline

Configuration

Vendors are set per-leg in the agent’s join request. Set vendor to "smallestai" on both the asr and tts blocks.

ASR (smallestai)

{
"asr": {
"vendor": "smallestai",
"language": "en",
"params": {
"api_key": "<smallest_ai_api_key>",
"url": "wss://api.smallest.ai/waves/v1/stt/live"
}
}
}
ParameterTypeDescription
api_keystringYour Smallest AI API key (required)
urlstringStreaming WebSocket endpoint. Use wss://api.us.smallest.ai/waves/v1/stt/live for East Asian languages (zh, yue, ja, ko), which are served from the US region only
languagestringLanguage code (e.g. "en", "hi"). Overrides the top-level ASR language field
sample_rateintegerInput audio sample rate in Hz. Defaults to 16000
encodingstringInput PCM encoding. Defaults to "linear16"
word_timestamps / sentence_timestampsstringInclude word- or sentence-level timing in results
diarizestringEnable speaker diarization
endpointingstringFinalize on trailing silence instead of a fixed timeout
eou_timeout_msstringSilence threshold (ms) before finalizing an utterance
punctuate / capitalizestringApply punctuation and capitalization to transcripts
itn_normalizestringConvert spoken numbers/dates to written form
keywordsstringBoost recognition, formatted as "term:weight,term:weight"
redact_pii / redact_pcistringMask personal/payment info in transcripts

Boolean-like parameters (diarize, endpointing, punctuate, etc.) are sent as the strings "true"/"false" - the Python, TypeScript, and Go SDKs convert native booleans automatically.

TTS (smallestai)

{
"tts": {
"vendor": "smallestai",
"params": {
"api_key": "<smallest_ai_api_key>",
"url": "https://api.smallest.ai/waves/v1/tts/live",
"model": "lightning_v3.1_pro",
"voice_id": "meher"
}
}
}
ParameterTypeDescription
api_keystringYour Smallest AI API key (required)
urlstringDefaults to https://api.smallest.ai/waves/v1/tts/live
modelstring"lightning_v3.1_pro" (premium English + Hindi voices) or "lightning_v3.1" (multilingual, cloning)
voice_idstringCatalog or cloned voice ID. Pro voices must be paired with lightning_v3.1_pro
languagestringLanguage code - see the model cards for supported lists per model
speednumberSpeech speed multiplier (0.5-2.0)
sample_rateintegerOutput audio sample rate in Hz
number_pronunciation_languagestringLanguage used to pronounce numbers when it differs from language
math_notationbooleanRead mathematical notation aloud instead of spelling out symbols
pronunciation_dictsarray[string]Pronunciation dictionary IDs to apply

Full Agent Config Example

A minimal join config wiring Smallest AI for both legs, with OpenAI as the LLM:

{
"name": "smallest-ai-agent",
"properties": {
"channel": "your-channel-name",
"token": "<agora_rtc_token>",
"agent_rtc_uid": "0",
"asr": {
"vendor": "smallestai",
"language": "en",
"params": {
"api_key": "<smallest_ai_api_key>",
"url": "wss://api.smallest.ai/waves/v1/stt/live"
}
},
"llm": {
"vendor": "openai",
"params": {
"api_key": "<openai_api_key>",
"model": "gpt-4o-mini"
}
},
"tts": {
"vendor": "smallestai",
"params": {
"api_key": "<smallest_ai_api_key>",
"url": "https://api.smallest.ai/waves/v1/tts/live",
"model": "lightning_v3.1_pro",
"voice_id": "meher"
}
}
}
}

Submit this as the body of the Conversational AI Engine join request described in Agora’s quickstart. Once the agent joins the channel, audio flows: RTC microphone track → Smallest AI Pulse (ASR) → LLM → Smallest AI Lightning (TTS) → RTC audio track.


Notes

  • Interruptions (barge-in) are handled by Conversational AI Engine itself: when a user speaks over the agent, the TTS leg is flushed and Lightning synthesis is cancelled mid-stream - no custom logic needed.
  • For the full, current parameter list on Agora’s side, see the ASR and TTS provider pages in Agora’s docs.
  • For issues or questions about the Smallest AI side of the integration, contact us on Discord.