Post call metrics

Post call metrics let you pull specific insights from conversations after they end.

View as Markdown

Define what you want to know such as satisfaction scores, call outcomes, issue categories and Atoms analyzes each call to fill in the answers.

Post-call metrics list

Post-call metrics dashboard

How It Works

  1. You define metrics: What questions do you want answered about each call?
  2. Call ends: Conversation completes normally
  3. AI analyzes: Atoms reviews the transcript against your metrics
  4. Data populated: Your metrics get filled in automatically
  5. Access anywhere: View in logs, receive via webhook, export

After the call, the transcript is sent to the post-call-analytics (PCA) model, which outputs a value for every metric you configured. Each output is validated (type check, enum membership, and so on) before it is stored and displayed. If any single metric fails validation, the whole post-call result for that call is rejected, none of the other metrics show up. So one bad field costs you the entire call’s analytics. Follow the authoring rules below to avoid this.

Creating a New Metric

Click the Add Metrics + button to open the configuration panel. You’ll see two options:

Disposition metrics

Create a new metric from scratch

Build a custom metric from scratch. Fill in the Identifier, Data Type, and Prompt. See details below.

Use Add Another + to create multiple metrics at once.

Don’t forget to hit Save in the Disposition tab once you’re done.

Configuring a Metric

Each metric needs three things:

FieldRequiredDescription
IdentifierYesUnique name for this metric. Lowercase, numbers, underscores only.
Data TypeYesWhat kind of value: String, Number, or Boolean
PromptYesThe question you want answered about the call

Identifier

This is the key used to reference the metric in exports, webhooks, and the API.

customer_satisfaction
call_outcome
follow_up_needed

Naming rules: Lowercase letters, numbers, and underscores only. No spaces or special characters.

Data Type

TypeUse forExample values
StringFree text, categories”resolved”, “escalated”, “billing issue”
BooleanYes/no questionstrue, false
IntegerWhole numbers, scores1, 5, 10
EnumFixed set of optionsOne of: “low”, “medium”, “high”
DatetimeDates and times”2024-01-15T10:30:00Z”

Prompt

This is the question the AI answers by analyzing the transcript. Be specific.

Good prompts:

  • “Did the agent acknowledge and respond to customer concerns effectively?”
  • “Rate customer satisfaction from 1 to 5 based on tone and words used.”
  • “What was the primary reason for this call? Options: billing, technical, account, other”

Vague prompts to avoid:

  • “Was it good?”
  • “Customer happy?”

Start with 3-5 metrics. Too many can slow analysis and clutter your data. Add more as you learn what insights matter most.

Authoring rules & common failure patterns

Get any of these wrong and the PCA model’s output fails validation for that call. When that happens no metric at all is stored for it, not just the one that broke. The five rules below cover most of the failure patterns we see in production.

1

Enum outputs must be fully enumerated

Anything you ask the model to output has to exist in the choices list. If the prompt can lead the model to emit a value (including null or "none") that isn’t in the enum, validation fails.

If you want “none” or “null” to be a valid outcome, add it as an explicit choice.

2

Never use skip / omit conditions on a metric

Conditions like “Omit this metric if callback_requested is No” cause the model to drop the field entirely, which fails validation and takes down the whole call’s analysis.

Keep the field always-present and use an explicit sentinel value instead. For example, “Return NA if callback is not requested” (with NA present in the enum).

3

Datatype must match the prompt

An integer-typed metric must never be instructed to output a string.

Common failure: an integer promised_amount field with a fallback like “output ‘Not Applicable’ if no amount was promised”. The string fails the integer type check.

Use a valid in-type sentinel instead (or a nullable/allowed value that matches the declared type).

4

Don’t create dependencies between metrics

One metric should not depend on another (e.g. lead_disposition depending on final_disposition, or a score depending on a disposition).

Metrics are extracted independently, not sequentially like variables in code, so cross-references are unreliable.

5

Fully specify any aggregation or weighting

If a rating is “a weighted score” across sub-metrics, define the weights explicitly (or state “simple average”).

Example Metrics

FieldValue
Identifiercall_outcome
Data TypeString
Prompt”What was the outcome of this call? Options: resolved, escalated, transferred, abandoned, callback_scheduled”

Configuring via API

Post-call metrics are set through the agent versioning flow: edit a draft’s config, publish it as a new version, and activate the version. There is no standalone post-call-analytics endpoint. Configuration lives on the agent’s active version.

Full flow

1import os
2import requests
3
4API_KEY = os.environ["SMALLEST_API_KEY"]
5AGENT_ID = "YOUR_AGENT_ID"
6BASE = "https://api.smallest.ai/atoms/v1"
7HEADERS = {"Authorization": f"Bearer {API_KEY}", "Content-Type": "application/json"}
8
9# 1. Get the currently active version
10agent = requests.get(f"{BASE}/agent/{AGENT_ID}", headers=HEADERS).json()
11active_version_id = agent["data"]["activeVersionId"]
12
13# 2. Create a draft from that version
14draft = requests.post(
15 f"{BASE}/agent/{AGENT_ID}/drafts",
16 headers=HEADERS,
17 json={"sourceVersionId": active_version_id, "draftName": "post-call-metrics-update"},
18).json()
19draft_id = draft["data"]["draftId"]
20
21# 3. Write your post-call metrics config onto the draft
22requests.patch(
23 f"{BASE}/agent/{AGENT_ID}/drafts/{draft_id}/config",
24 headers=HEADERS,
25 json={
26 "postCallAnalyticsConfig": {
27 "dispositionMetrics": [
28 {
29 "identifier": "call_resolved",
30 "dispositionMetricPrompt": "Was the customer issue resolved by the end of the call?",
31 "dispositionMetricType": "BOOLEAN",
32 },
33 {
34 "identifier": "satisfaction_score",
35 "dispositionMetricPrompt": "Rate the customer satisfaction from 1 to 5",
36 "dispositionMetricType": "INTEGER",
37 },
38 {
39 "identifier": "call_outcome",
40 "dispositionMetricPrompt": "What was the outcome of the call?",
41 "dispositionMetricType": "ENUM",
42 "choices": ["resolved", "escalated", "callback_scheduled", "no_action"],
43 },
44 {
45 "identifier": "summary",
46 "dispositionMetricPrompt": "Provide a brief summary of the call",
47 "dispositionMetricType": "STRING",
48 },
49 ],
50 "useInternalAnalyticsModel": True,
51 "useReasoningModel": False,
52 }
53 },
54).raise_for_status()
55
56# 4. Publish the draft as a new version and activate it
57requests.post(
58 f"{BASE}/agent/{AGENT_ID}/drafts/{draft_id}/publish",
59 headers=HEADERS,
60 json={"label": "Added post-call metrics", "activate": True},
61).raise_for_status()

Disposition metric schema

Each entry in dispositionMetrics takes four fields:

FieldRequiredDescription
identifierYesMachine name. Lowercase letters, digits, and underscores only.
dispositionMetricPromptYesThe question evaluated against the transcript.
dispositionMetricTypeYesOne of STRING, BOOLEAN, INTEGER, ENUM, DATETIME.
choicesOnly when type is ENUMAllowed values.

Two analytics flags apply globally:

FieldDefaultDescription
useInternalAnalyticsModeltrueUse the internal analytics model. When false, the agent’s own LLM evaluates metrics.
useReasoningModelfalseRoute evaluation through the reasoning model (higher quality, higher latency/cost).

Reading current metrics

$curl -H "Authorization: Bearer $SMALLEST_API_KEY" \
> "https://api.smallest.ai/atoms/v1/agent/$AGENT_ID" \
> | jq '.data._resolvedConfig.postCallAnalyticsConfig'

The active version’s metrics are surfaced under _resolvedConfig.postCallAnalyticsConfig. Each entry round-trips exactly as sent: identifier, dispositionMetricPrompt, dispositionMetricType, and choices (when the type is ENUM).

FAQ

Post-call analytics runs after the call ends, so it takes a moment to populate. If a call finished cleanly but nothing shows in the Conversations log after a minute, it is almost always a validation failure on one of the metrics you configured. Check the authoring rules. A single mis-authored field (wrong datatype, missing enum value, skip condition, etc.) rejects the entire call’s analytics.

Almost always one of the authoring patterns above. The most common causes:

  • An enum that doesn’t cover every value the model can produce.
  • A metric prompt that asks the model to output a value in a type other than what’s declared (e.g. a string fallback on an integer field).
  • A “weighted score” or “aggregate” metric with no weights defined, so the model invents them.

Review the metric definition against the authoring rules.

No. Each metric is extracted independently, not sequentially. If lead_disposition depends on the value the model returned for final_disposition, the reference is unreliable. Encode any cross-metric logic inside the individual metric’s prompt (using the transcript as the source of truth), not by referencing sibling metrics.