Every QA program starts roughly the same way. A reviewer grades 20 of 1,000 tickets, and the number goes onto a slide: Quality: 92%. The other 98% were never scored. Sampling was the only option when a person had to read each one. AI changes the constraint.
Key takeaways
- A failure that happens 5 times in 10,000 conversations has roughly a 90% chance of never appearing in a 2% sample.
- Automate the checkable criteria, let AI score interpretive ones under human calibration, and keep accountability human.
- Reviewers move to calibration, exceptions, coaching and root cause.
- Scored at 100%, QA becomes a dataset you can combine with contact drivers and resolution.
What does a 2% QA sample actually miss?
Often, every instance of a rare failure.
Suppose your team handles 10,000 conversations and randomly reviews 2%, or 200. Here’s the approximate probability that those 200 reviews contain zero examples of a failure at different rates:
| Failure rate | Instances in 10,000 | Approx. chance a 200-conversation sample misses all of them |
|---|---|---|
| 5% | 500 | <0.01% |
| 1% | 100 | ~13% |
| 0.5% | 50 | ~36% |
| 0.2% | 20 | ~67% |
| 0.1% | 10 | ~82% |
| 0.05% | 5 | ~90% |
These are approximate probabilities under random sampling, not observed customer data: the chance of missing every instance is about 0.98 raised to the number of instances.
If a serious failure occurs only five times in 10,000 conversations, a 2% random review has roughly a 90% chance of missing every instance. Those rare events are often the failures you care most about: a skipped disclosure, a promise policy doesn’t allow, a misapplied refund policy, or an AI agent giving a wrong answer.
The 5 failures sit outside the sample.
chance none of the 5 serious failures is ever reviewed
Every eligible conversation scored against the rubric.
serious failures scored, each with the conversation behind it
Klaviyo, handling over 600,000 incidents a year, had been reviewing under 2% of them (Klaviyo story).
What does “100% QA” actually mean?
It means every eligible conversation in the population you’ve defined and connected is evaluated. “100% of Zendesk support tickets from September 1–30” is a defensible measurement. “100% of customer conversations” is only accurate if every relevant source is actually represented.
What changes when you QA every support conversation?
QA becomes a dataset.
Every eligible conversation can carry structured fields: overall QA score, policy adherence, resolution, required disclosure completed, correct next step, escalation quality, customer effort, agent, contact driver and root cause.
Instead of “Our quality score was 88% last month,” you can ask: Which contact drivers produced our lowest QA scores? Which agents need coaching specifically on escalation quality?
The score is another structured layer on the conversation dataset.
Which QA criteria should AI score?
Not every scorecard line should be treated the same way.
Divide the rubric into three tiers:
| Tier | Criteria | Who owns the score |
|---|---|---|
| Automate the checkable | Required disclosure, verification step, correct policy, correct next step, resolution confirmed, escalation path, prohibited language | AI, with the evidence |
| AI scores, humans calibrate | Problem understood, clear explanation, customer effort, de-escalation, concern acknowledged | AI for measurement and prioritization; a human confirms before a score drives coaching or performance |
| Keep accountability human | Formal performance actions, disputed scores, novel situations, rubric calibration, ambiguous edge cases | A person |
Interpretive criteria need stronger calibration because reasonable reviewers disagree about what “good” looks like. AI can own measurement without owning accountability. (More in the complete guide to conversation analytics.)
How do you know whether AI QA scores are trustworthy?
Calibrate before rolling scores out.
Have experienced reviewers and the AI independently score a shared set of conversations with your existing rubric, then compare results criterion by criterion, not just overall.
When they disagree, inspect why. Sometimes the model is wrong; sometimes two humans disagree too, which says something about the rubric. “Demonstrated empathy” leaves enormous room for interpretation. “Did the agent acknowledge the customer’s frustration before the next troubleshooting step?” is easier to apply consistently. Calibration forces you to make the definition of quality explicit.
Where do human QA reviewers go when AI scores every conversation?
Nowhere. Their work gets more valuable:
- Calibration: test whether AI and reviewers interpret the rubric consistently.
- Exceptions: investigate the lowest-scoring, highest-risk or most unusual conversations.
- Coaching: instead of “Your QA score is 84%,” show the two behaviors behind lower scores, with examples from the agent’s own conversations.
- Root cause: the same failure 400 times is probably not 400 coaching problems; the policy, macro or help article may be wrong.
Why does traceability matter for automated QA scores?
Every automated QA score should carry evidence.
If the system scores Required disclosure: FAIL, a reviewer should be able to open the conversation and see why. People won’t trust a score they can’t challenge, and evidence shows whether the model, the rubric or the source data was at fault.
“Can you just refund it to a different card?”
“Done, it’s on its way.”
How does Rippit handle QA?
Rippit treats QA as one type of analysis on the same conversation data used for customer insights, churn signals and chatbot monitoring. An admin connects your customer conversation hubs to Rippit in one click—including Zendesk, Intercom, Gong and more—and a team defines a QA agent in plain language:
Through MCP, the agent can also read, write and update data across your broader stack, including Salesforce, HubSpot, Slack, Notion, Guru, Jira, Snowflake, and more.
Those outputs become structured fields you can combine with others. See how it works, ready-made skills, and why Rippit reads 100% of conversations. Brex’s public story describes rebuilding its QA playbook from the ground up.
Comparing QA tools? See how Rippit stacks up against Zendesk QA, Level AI and Solidroad.
FAQ: QA for 100% of support conversations
Does full-coverage QA mean every score is automatically correct?
No. Full coverage solves the coverage problem; calibration solves measurement; humans handle accountability. A bad rubric applied to 100% of conversations is still a bad rubric.
Can AI QA scores be used for performance reviews?
Use caution. Before an AI-derived criterion feeds a formal employment decision, validate it against experienced reviewers and set up a process for disputed scores.
Do you still need manual QA?
Yes, but much less random manual scoring. Humans stay on calibration, coaching, appeals and improving the rubric.
Is this different from what MaestroQA did?
MaestroQA is Rippit’s former name, not a different product; QA customers use the same platform. QA is now one of several agents teams build on the same data.
How long does it take to get the first scores?
Connect your customer conversation hubs to Rippit in one click—including Zendesk, Intercom, Gong and more—then describe the agent in plain language; it scores on a schedule. Through MCP, the agent can also read, write and update data across your broader stack, including Salesforce, HubSpot, Slack, Notion, Guru, Jira, Snowflake, and more. Budget more time for calibration than setup.
.webp)