Key takeaways
- An escalation audit answers a different question than a containment dashboard: not how often the bot handed off, but why it handed off and whether it should have.
- Classifying every escalation into five root causes (retrieval failure, incorrect or stale knowledge, knowledge gap, capability or policy limit, and correct handoff or customer exit) turns one number into five fix lists with different owners.
- Handoff quality is measured separately: a correct escalation can still be a bad customer experience if context is lost or the customer has to repeat themselves.
- An escalation is clearly avoidable when the bot already had the knowledge it needed and failed to use it.
- Sampling a small share of escalations will reliably find your most common failure and reliably miss the rare, expensive one.
Why isn’t escalation rate enough?
Most teams running an AI support agent monitor some version of containment rate or escalation rate. Those metrics tell you what happened. They don’t tell you why.
Consider three conversations:
- A customer immediately types, “I want to talk to a person.”
- The bot can’t find an existing refund policy and escalates.
- The bot confidently gives the wrong refund policy, the customer pushes back, and only then does it escalate.
All three count as escalations. But the first may be correct behavior. The second may be a retrieval problem. The third is a quality incident. Combine them into one escalation rate and you’ve created a metric without a remediation plan.
A useful escalation audit therefore produces two separate outputs:
Why did the bot hand off?
Once it handed off, how well did it do it?
Here’s how to build both.
Step 1: Define every escalation you want to audit
Audit every conversation where the bot handed off, a human took over, or the customer asked for a person or gave up, not a sample.
Include conversations where:
- The bot explicitly handed off to a human.
- A human joined after the bot attempted to resolve the issue.
- The customer explicitly requested a human.
- The customer abandoned after an unsuccessful bot exchange where escalation should reasonably have occurred.
Then analyze the entire population. Sampling made sense when reviewing conversations was expensive. If someone had to manually read 20,000 transcripts, reviewing 2% was often the only practical option.
That constraint is disappearing. If an AI agent can evaluate every connected conversation, there is little reason to use a small sample as your primary measurement system.
That’s especially important for AI-agent monitoring because the highest-volume problem isn’t necessarily the highest-risk problem. Incorrect pricing, bad policy guidance, mishandled sensitive situations, or other serious failures may occur infrequently enough to disappear from a small sample.
With Rippit, teams can create an agent that evaluates the complete set of connected bot conversations using criteria written in plain language. Connect your customer conversation hubs to Rippit in one click—including Zendesk, Intercom, Gong and more—then tell the agent what to check across 100% of your connected conversations. Through MCP, the agent can also read, write and update data across your broader stack, including Salesforce, HubSpot, Slack, Notion, Guru, Jira, Snowflake, and more. That means an audit can check what the bot said against your help center, your CRM record or your verified knowledge base, not just the transcript.
The result is a measurement of the actual population rather than an extrapolation from a handful of transcripts. For the wider monitoring approach, see The Complete Playbook for AI Agent Monitoring.
Step 2: Classify every escalation by root cause
Give every escalation exactly one primary root cause from the five codes below, instead of counting it in a single “escalated” bucket.
| Code | Root cause | What it looks like | Primary owner |
|---|---|---|---|
| RC-01 | Retrieval failure | The correct knowledge existed, but the bot failed to find or use it | AI / retrieval owner |
| RC-02 | Incorrect or stale knowledge | The bot found information, but the source itself was wrong, incomplete, or outdated | Knowledge / content |
| RC-03 | Knowledge gap | The request was in scope, but no approved answer existed | Knowledge / content |
| RC-04 | Capability or policy limit | The bot understood the request but wasn’t allowed or able to complete the required action | Product / operations |
| RC-05 | Correct handoff or customer exit | The request required a human by design, or the customer explicitly requested one | CX / AI-agent owner |
Assign one primary root cause to each escalation. If several things went wrong, choose the earliest issue that, had it been fixed, would have changed the outcome.
The distribution is more useful than the overall escalation rate because each bucket creates a different action:
- A spike in RC-01 means you investigate retrieval.
- A spike in RC-02 means you fix source content.
- RC-03 identifies missing knowledge.
- RC-04 tells you where additional integrations, permissions, or automation could expand containment.
- And RC-05 may not need fixing at all.
That last point matters. The goal should not be zero escalations. The goal should be zero bad escalations.
Step 3: Score the quality of the bot-to-human handoff
A necessary escalation can still be a terrible customer experience, so score handoff quality separately from root cause.
Here’s a five-dimension rubric. Score each dimension from 0–2 for a total possible score of 10.
| Dimension | 0 | 1 | 2 |
|---|---|---|---|
| EQ-1: Timeliness | Bot loops three or more times before escalating | One unnecessary attempt | Escalates at first clear signal it can’t resolve |
| EQ-2: Context transfer | Human starts cold | Partial context transferred | Human receives full issue summary and prior attempts |
| EQ-3: Customer repetition | Customer repeats entire problem | Customer repeats some information | Customer repeats nothing |
| EQ-4: Expectation setting | No explanation of what happens next | Handoff acknowledged but unclear | Clear next step or wait expectation |
| EQ-5: Outcome integrity | Bot leaves incorrect or invented information on record | Answer is incomplete or ambiguous | No incorrect information left behind |
The total score is useful, but EQ-5 deserves special treatment. If the bot invented a policy, quoted the wrong price, gave unsafe instructions, or otherwise left the customer with incorrect information before handing off, a decent context transfer shouldn’t hide that.
Treat an EQ-5 score of zero as a separate quality incident.
Zendesk’s own documentation describes evaluating AI agents with scorecards, either manually or automatically, and monitoring escalations through a BotQA dashboard (Evaluating the performance of AI agents using Zendesk QA). Check whether your setup can apply a custom rubric like this one across bot conversations in channels outside that helpdesk, and whether the same rubric can run against human conversations for comparison. For a side-by-side view, see Rippit vs. Zendesk.
Step 4: Separate avoidable from correct escalations
An escalation is avoidable when the bot could reasonably have resolved the request using the knowledge, permissions, and capabilities it was intended to have at that time.
That answers the question executives actually care about: how much of this could we have prevented? It generally makes RC-01 the clearest avoidable category. Parts of RC-02 and RC-03 may also be preventable, but they represent different problems: the bot cannot correctly use knowledge that the business itself hasn’t maintained or created.
So don’t stop at a single “avoidable” percentage. Report the underlying causes with it. Using the worked example below (4,000 bot conversations):
of bot conversations escalated
bot failed to use knowledge it already had (RC-01)
knowledge was stale or missing (RC-02, RC-03)
policy, capability or correct handoffs (RC-04, RC-05)
Now you have something you can act on. “Escalation is 31%” starts a debate. “7.8% of conversations escalated even though we already had the information needed to answer them” creates a backlog.
Step 5: Save audit results as structured data
Save each audit result as a structured field on the conversation, so the next question doesn’t require a new analysis.
This is where escalation auditing becomes much more powerful than a one-time AI prompt. For every bot conversation, you can create structured fields such as:
- Escalated: yes/no
- Escalation root cause
- Avoidable: yes/no
- Handoff quality score
- Context transferred: yes/no
- Customer repeated themselves: yes/no
- Incorrect information before escalation: yes/no
- Knowledge article involved
- Final resolution
- Customer sentiment
Now those results aren’t trapped inside an AI-generated summary. They become part of the conversation dataset. That means you can ask:
“Are customers with poor handoff scores more likely to contact us again within seven days?”
“Which AI-agent intents have the highest rate of RC-01 retrieval failures?”
“Did our RC-02 rate fall after last week’s help-center update?”
You don’t have to reconstruct the original analysis every time. The output of the AI becomes reusable data.
With Rippit, a non-technical user can define these fields and agents in natural language, run them across connected conversations, and schedule the analysis to update as new conversations arrive, with results delivered to email or Slack. Through MCP, the same agent can push findings into the tools where the fixes happen, such as an issue in Jira or Linear for each knowledge gap. Ready-made workflows for chatbot analysis and escalation review are documented in Rippit skills. For the broader method, see The Complete Guide to Conversation Analytics.
What does a completed escalation audit look like?
A completed audit turns one escalation rate into a ranked list of causes, each with an owner and a fix.
The example below is a fictional composite with round numbers for illustration; these are not measured results from any customer. Consider a fictional support team with 4,000 bot conversations in one week. Of those, 1,240 escalated, for an overall escalation rate of 31%. After classifying the conversations:
| Root cause | Escalations | Share |
|---|---|---|
| RC-01 Retrieval failure | 310 | 25% |
| RC-02 Incorrect/stale knowledge | 124 | 10% |
| RC-03 Knowledge gap | 185 | 15% |
| RC-04 Capability/policy limit | 385 | 31% |
| RC-05 Correct handoff/customer exit | 236 | 19% |
Now the 31% escalation rate tells a story. The team discovers that retrieval failures account for 310 conversations. More importantly, six knowledge articles are associated with most of those failures.
Separately, handoff scoring shows that context transfer and customer repetition are the weakest dimensions. Those are two different problems with two different fixes: one team improves retrieval; another fixes the handoff payload so human agents receive the bot’s context.
That’s what an escalation audit should produce: not another dashboard, but a prioritized list of things to fix.
Who should use Rippit to audit chatbot escalations?
With Rippit, anyone can build an agent that reviews AI agent and chatbot conversations alongside human ones, so teams can see where the bot resolves, escalates or gets answers wrong. Rippit analyzes each conversation once when it arrives, so an escalation audit runs across the full set rather than a sample, and every number traces back to the conversations behind it. For a customer example, see the Klaviyo customer story.
- Teams running an AI support agent at enough volume that manual escalation review has stopped being feasible
- CX and support ops leaders who need to show which escalations were avoidable
- Regulated teams that can’t rely on sampled review
- Teams whose bot handles a handful of conversations a day, where reading all of them by hand is faster
- Teams that want to instrument model-level traces and latency, which is an LLM observability problem, not a conversation analysis one
- Teams looking for a tool that rewrites bot flows for them: this audit tells you what to change, it doesn’t change it
If you run a vendor AI agent, see how Rippit compares with Intercom Fin, Sierra and Decagon.
What is Rippit?
Rippit is a conversation data platform that lets anyone build AI agents using all of their conversation data. You don’t have to be an engineer. Connect your customer conversation hubs in one click—including Zendesk, Intercom, Gong and more—describe what your agent should do in plain language, and it reads 100% of your connected conversations and delivers results on a schedule. Through MCP, the agent can also read, write and update data across your broader stack, including Salesforce, HubSpot, Slack, Notion, Guru, Jira, Snowflake, and more. QA, insights, churn, chatbot monitoring and compliance are all kinds of agents people build. See how Rippit works.
FAQ: auditing AI chatbot escalations
How often should you audit AI chatbot escalations?
For a newly deployed or frequently changing AI agent, run the analysis at least weekly. Once performance stabilizes, the cadence can move to monthly, but severe quality conditions such as incorrect information should be monitored continuously or as frequently as your operation requires. Apply the same definitions consistently over time so you can see whether the underlying failure modes are actually improving. An agent built in Rippit can run the audit on a schedule and deliver results to email or Slack.
Is escalation rate a bad metric?
It is an incomplete one. Escalation rate mixes correct handoffs with failures, so it can fall while customer experience gets worse, for example if the bot stops escalating cases it should escalate. Track avoidable escalation rate and handoff quality alongside it.
How is this different from LLM observability?
LLM observability and conversation analysis answer different questions. Observability tools inspect the system: model calls, traces, latency, token usage, tool calls and errors. An escalation audit inspects the customer interaction: did the customer get the right answer, why couldn’t the bot resolve the request, was escalation appropriate, was context preserved, did the customer have to repeat themselves, and what ultimately happened. You may need both. A technically successful model call can still produce a terrible customer experience.
Can you audit chatbot escalations without an engineer?
Yes. The root-cause taxonomy and QA rubric above are written as plain-language judgments. In Rippit, teams can turn judgments like these into agents and structured fields without writing SQL or building their own data pipeline. Rippit ingests the underlying conversation data, runs the analysis across connected conversations, and maintains the results as part of the conversation dataset. Rippit hosts its MCP server itself, so there is nothing to deploy (MCP server overview). That makes the audit repeatable instead of a one-time project.
What if our chatbot isn’t in Intercom or Zendesk?
Run the playbook anyway: the taxonomy and rubric are vendor-neutral and work against any transcript your team can export. Rippit connects your customer conversation hubs in one click—including Zendesk, Intercom, Gong and more—and through MCP the agent can also read, write and update data across your broader stack, including Salesforce, HubSpot, Slack, Notion, Guru, Jira, Snowflake, and more. New integrations ship regularly; see the current list on the integrations page and in the supported integrations docs.
The bottom line
Don’t optimize your AI agent for the lowest possible escalation rate. Optimize for correct resolution when the bot should resolve, and a clean handoff when it shouldn’t.
- Measure why every escalation happened.
- Measure the quality of the handoff separately.
- Turn those measurements into persistent data.
- Then fix the failure modes that actually matter.
That’s how escalation rate becomes an operating system instead of a vanity metric.
Zendesk Help Center · Evaluating the performance of AI agents using Zendesk QA
Rippit · Klaviyo customer story
Rippit · Integrations
Rippit Docs · Supported integrations
Rippit Docs · MCP server overview
Rippit Docs · Skills
.webp)