ANSWER

Why Chatbot Containment Rate Misleads, and What to Measure Instead

Containment tells you whether a chatbot conversation ended without a human. It does not tell you whether the customer’s problem was solved.

Try Rippit FreeRequest Demo

Containment rate tells you whether a chatbot conversation ended without a human. It does not tell you whether the customer’s problem was solved. A customer whose question was answered correctly can count as contained. So can a customer who received a wrong answer, got frustrated, closed the chat, and emailed support the next morning. Only one was a success.

Containment measures the absence of a human. It does not measure the presence of an answer. Measure the customer outcome instead: confirmed resolution, unresolved abandonment, repeat contact, handoff quality, and wrong answers.

With Rippit (formerly MaestroQA), anyone can build an agent that reviews AI agent and chatbot conversations alongside human ones, so teams can see where the bot resolves, escalates or gets answers wrong.

Key takeaways

  • Containment measures the absence of a human, not the presence of an answer, so an abandoned chat and a solved problem score identically.
  • Replace one number with five: confirmed resolution, unresolved abandonment, repeat contact rate, handoff quality and wrong-answer rate.
  • Confirmed resolution needs a window and a repeat-contact check; without both, you are measuring session endings.

What does chatbot containment rate actually measure?

Containment asks one question: did this interaction end without escalating to a human agent?

The problem comes when containment becomes a proxy for resolution.

Consider three conversations:

  • Conversation A: The bot gives the correct answer and the customer leaves.
  • Conversation B: The bot confidently gives the wrong answer and the customer accepts it.
  • Conversation C: The bot repeats the same troubleshooting steps until the customer stops responding.

All three can appear contained. Good AI-agent monitoring asks what actually happened to the customer’s issue.

All three conversations count as contained; only one solved the customer’s problem. Illustrative example; numbers are not real data.

What should you measure instead of chatbot containment? The Resolution Ledger

No single metric replaces it. You need five that describe different failure modes.

The Resolution Ledger: five metrics to report alongside containment
MetricWhat it measuresMeasurement windowFailure signal
Confirmed resolutionEvidence that the customer’s issue was resolved, with no same-issue repeat contactSession + the repeat-contact windowIssue appears unresolved or customer returns with the same problem, in any channel
Unresolved abandonmentCustomer stops engaging before there is evidence of resolution or successful handoffSessionConversation ends without evidence the issue was solved
Repeat contact rateCustomer contacts support again about the same underlying issueOften 72 hours; adapt to your businessSame customer, same issue, another conversation in any channel
Handoff qualityWhether escalation happened at the right time and context carried to the humanHandoffCustomer has to repeat information already given to the bot
Wrong-answer rateBot gives information that conflicts with the applicable policy, documentation, or known factsSessionMaterially incorrect or unsupported answer

Set the window by product (a weekly-use B2B tool needs longer than checkout). The important part is having a window at all.

A resolution metric without a repeat-contact window is just a session-end counter with a better name.

Because customers don’t always come back, confirmed resolution should combine signals: the requested action was completed, the customer confirmed the fix, and no same-issue contact followed during the window.

Why measure the customer’s issue, not just the chatbot session?

Containment is session-centric. Customers are not.

A customer chats with the bot about a failed payment on Monday, chats again on Tuesday, and emails support on Wednesday, where a human finally fixes it. A dashboard sees three interactions; the customer had one problem that took three conversations to resolve.

The better data model is conversation → issue → outcome → repeat contact, rather than session → contained / not contained. Because customers rephrase (“My payment won’t go through,” then “Still can’t complete checkout”), matching ticket tags undercounts; analyze each conversation’s contact driver to decide whether a later one is the same problem.

What can a 40% containment rate look like after you measure outcomes?

Consider a hypothetical example.

An AI support agent handles 1,000 conversations in a month, and its dashboard reports 40% containment: 400 conversations ended without a human. Inspect those 400, counting each conversation once:

  • 80 ended in unresolved abandonment.
  • 55 were followed by another contact about the same issue.
  • 15 contained a materially incorrect answer.
  • 250 showed evidence of resolution and survived the repeat-contact check.

The headline metric was containment: 40% of all conversations. The more meaningful one was confirmed resolution: 25% of all conversations (250 of 1,000), which is 62.5% of the contained ones. A containment rate tells you where the conversation ended. An outcome analysis tells you what happened to the problem.

Inside the 400 contained conversations (each counted once)
OutcomeConversationsShare of contained
Confirmed resolution25062.5%
Unresolved abandonment8020%
Repeat contact, same issue5513.75%
Materially incorrect answer153.75%
A 40% containment rate becomes 25% confirmed resolution once each contained conversation is checked for its outcome. Illustrative example; numbers are not real data.

What makes a chatbot handoff good, and how do you catch wrong answers?

Escalation isn’t necessarily a failure; the question is whether the handoff was appropriate and effective: did the bot escalate at the right point, collect useful information first and pass that context on, so the customer didn’t repeat themselves and the human resolved the issue? A bot that escalates 30% of conversations cleanly may serve customers better than one that “contains” 90% by refusing to give up.

Wrong answers are dangerous because they look like successful containment: the customer leaves and the dashboard records a win. Evaluate what the bot said against the relevant source of truth, such as current policy, help-center documentation or product behavior, and classify each answer as correct, incomplete, unsupported, or wrong.

How does Rippit measure AI-agent outcomes?

Rippit turns customer conversations into structured, queryable data and lets teams build AI agents on top of it without engineering.

Define an analysis in plain language: Review every chatbot conversation. Determine whether the customer’s underlying issue was resolved, whether they abandoned first, whether the bot gave an incorrect answer, and whether any handoff preserved the context they already provided. Rippit applies those classifications across every connected conversation and keeps the results, so you can ask which contact drivers have the highest unresolved-abandonment rate. (How it works.)

Connect your customer conversation hubs to Rippit in one click—including Zendesk, Intercom, Gong and more—then define agents in plain language and run them repeatedly across conversation data. Through MCP, the agent can also read, write and update data across your broader stack, including Salesforce, HubSpot, Slack, Notion, Guru, Jira, Snowflake, and more. Results can be read inside an AI assistant through Rippit’s hosted MCP server: an admin adds one connector URL, there is nothing to deploy, and each user sees only what they can already see in Rippit.

Klaviyo’s public customer story describes running AI across 100% of bot conversations to surface silent failures and find where humans were succeeding where the bot wasn’t, then feeding those patterns back into the bot’s training, and routing an immediate handoff for intents the bot consistently struggled with, prioritizing speed to resolution over containment.

Evaluating AI support agents? See how Rippit compares with Intercom Fin, Sierra and Decagon.

FAQ: chatbot containment rate

Is containment rate still useful?

Yes, as an operational metric. It tells you how often the bot prevents a human interaction, which matters for staffing and cost modeling. Report it alongside the five Resolution Ledger metrics, so you can tell whether the bot reduced human work because it solved the problem or because the customer stopped talking.

What’s a good chatbot resolution rate?

There isn’t one universal number worth optimizing against, because definitions and windows differ. A more useful benchmark is your own AI agent over time, segmented by contact driver, including where humans materially outperform the bot.

How is this different from LLM observability?

LLM observability monitors the model’s system: latency, tokens, errors, traces. Conversation monitoring measures what happened to the customer: was the answer correct, was the problem solved, did the customer come back? You need both. A perfectly healthy model endpoint can still give customers bad answers.

Can I compare two AI-agent vendors this way?

Yes. Vendors define containment differently, so their dashboard percentages don’t compare. Apply the same definitions to both sets of conversations and you’re measuring the outcome using one standard. For the wider method, see the complete guide to conversation analytics.

What’s the bottom line on chatbot containment rate?

Containment answers “did a human get involved?” The question you care about is “did the customer’s problem get solved?” Keep containment for planning, but measure AI-agent quality by the outcome. Because a customer closing the window isn’t a resolution. Sometimes they just gave up.

Where conversations become

insights

actionable data

business intelligence

enterprise visibility

insights

I fear not the man who has practiced 10,000 kicks once, but I fear the man who has practiced one kick 10,000 times
Peloton
legal zoom