Skip to navigation

MCP: prompt-cache hit rate and per-turn LLM timings in debug_call, new get_latency_summary tool

debug_call now returns the LLM-side numbers for a call alongside the caller-perceived latency it already reported: usage (prompt, completion and cached tokens, LLM call count, and prompt-cache hit percentage), turns (per-turn LLM time to first token, generation time and total turn time) and toolCalls (each tool’s execution time and context tokens). Before this, those numbers were only recoverable by parsing the raw event timeline.

New get_latency_summary tool: caller-perceived latency KPIs for the org over a date range (average, p50, p95, p99), a daily trend, and average/p95 per pipeline stage, filterable to one agent.

The same usage, turns and toolCalls fields are documented on GET /conversation/{id} in the API reference. They have been returned in production; this change documents them.

→ MCP tool reference