ANSWER

How to QA 100% of Support Conversations Instead of 2%

Stop using humans as the sampling mechanism: let AI score every eligible conversation against your existing rubric, then point your reviewers at calibration, exceptions and coaching.

Try Rippit FreeRequest Demo
To QA 100% of support conversations, stop using humans as the sampling mechanism: let AI score every eligible conversation against your existing rubric, then point your reviewers at calibration, exceptions and coaching. With Rippit (formerly MaestroQA), anyone can build a QA agent that scores 100% of customer conversations and connects quality results to customer insights.

Every QA program starts roughly the same way. A reviewer grades 20 of 1,000 tickets, and the number goes onto a slide: Quality: 92%. The other 98% were never scored. Sampling was the only option when a person had to read each one. AI changes the constraint.

Key takeaways

  • A failure that happens 5 times in 10,000 conversations has roughly a 90% chance of never appearing in a 2% sample.
  • Automate the checkable criteria, let AI score interpretive ones under human calibration, and keep accountability human.
  • Reviewers move to calibration, exceptions, coaching and root cause.
  • Scored at 100%, QA becomes a dataset you can combine with contact drivers and resolution.

What does a 2% QA sample actually miss?

Often, every instance of a rare failure.

Suppose your team handles 10,000 conversations and randomly reviews 2%, or 200. Here’s the approximate probability that those 200 reviews contain zero examples of a failure at different rates:

Chance a 2% random sample (200 of 10,000 conversations) misses every instance of a failure
Failure rateInstances in 10,000Approx. chance a 200-conversation sample misses all of them
5%500<0.01%
1%100~13%
0.5%50~36%
0.2%20~67%
0.1%10~82%
0.05%5~90%

These are approximate probabilities under random sampling, not observed customer data: the chance of missing every instance is about 0.98 raised to the number of instances.

If a serious failure occurs only five times in 10,000 conversations, a 2% random review has roughly a 90% chance of missing every instance. Those rare events are often the failures you care most about: a skipped disclosure, a promise policy doesn’t allow, a misapplied refund policy, or an AI agent giving a wrong answer.

The same 10,000 conversations with 5 rare failures: a 2% sample usually misses all of them; scoring every conversation surfaces each one. Illustrative example; numbers are not real data.

Klaviyo, handling over 600,000 incidents a year, had been reviewing under 2% of them (Klaviyo story).

Use samples to calibrate the scorecard. Use the full dataset to measure quality.

What does “100% QA” actually mean?

It means every eligible conversation in the population you’ve defined and connected is evaluated. “100% of Zendesk support tickets from September 1–30” is a defensible measurement. “100% of customer conversations” is only accurate if every relevant source is actually represented.

What changes when you QA every support conversation?

QA becomes a dataset.

Every eligible conversation can carry structured fields: overall QA score, policy adherence, resolution, required disclosure completed, correct next step, escalation quality, customer effort, agent, contact driver and root cause.

Instead of “Our quality score was 88% last month,” you can ask: Which contact drivers produced our lowest QA scores? Which agents need coaching specifically on escalation quality?

The score is another structured layer on the conversation dataset.

Which QA criteria should AI score?

Not every scorecard line should be treated the same way.

Divide the rubric into three tiers:

Three tiers of QA criteria and who owns each score
TierCriteriaWho owns the score
Automate the checkableRequired disclosure, verification step, correct policy, correct next step, resolution confirmed, escalation path, prohibited languageAI, with the evidence
AI scores, humans calibrateProblem understood, clear explanation, customer effort, de-escalation, concern acknowledgedAI for measurement and prioritization; a human confirms before a score drives coaching or performance
Keep accountability humanFormal performance actions, disputed scores, novel situations, rubric calibration, ambiguous edge casesA person

Interpretive criteria need stronger calibration because reasonable reviewers disagree about what “good” looks like. AI can own measurement without owning accountability. (More in the complete guide to conversation analytics.)

How do you know whether AI QA scores are trustworthy?

Calibrate before rolling scores out.

Have experienced reviewers and the AI independently score a shared set of conversations with your existing rubric, then compare results criterion by criterion, not just overall.

When they disagree, inspect why. Sometimes the model is wrong; sometimes two humans disagree too, which says something about the rubric. “Demonstrated empathy” leaves enormous room for interpretation. “Did the agent acknowledge the customer’s frustration before the next troubleshooting step?” is easier to apply consistently. Calibration forces you to make the definition of quality explicit.

Where do human QA reviewers go when AI scores every conversation?

Nowhere. Their work gets more valuable:

  • Calibration: test whether AI and reviewers interpret the rubric consistently.
  • Exceptions: investigate the lowest-scoring, highest-risk or most unusual conversations.
  • Coaching: instead of “Your QA score is 84%,” show the two behaviors behind lower scores, with examples from the agent’s own conversations.
  • Root cause: the same failure 400 times is probably not 400 coaching problems; the policy, macro or help article may be wrong.
QA teams stop being a sampling function and start being a quality function.

Why does traceability matter for automated QA scores?

Every automated QA score should carry evidence.

If the system scores Required disclosure: FAIL, a reviewer should be able to open the conversation and see why. People won’t trust a score they can’t challenge, and evidence shows whether the model, the rubric or the source data was at fault.

One scored conversation: each criterion becomes a structured field, and a failing score links to the exact lines behind it so a reviewer can verify or dispute it. Illustrative example; numbers and quotes are not real data.

How does Rippit handle QA?

Rippit treats QA as one type of analysis on the same conversation data used for customer insights, churn signals and chatbot monitoring. An admin connects your customer conversation hubs to Rippit in one click—including Zendesk, Intercom, Gong and more—and a team defines a QA agent in plain language:

“Evaluate every support conversation against these eight criteria. For each criterion, return the score and the evidence supporting it. Flag conversations that fail any critical policy requirement and summarize the most common failure patterns each week.”

Through MCP, the agent can also read, write and update data across your broader stack, including Salesforce, HubSpot, Slack, Notion, Guru, Jira, Snowflake, and more.

Those outputs become structured fields you can combine with others. See how it works, ready-made skills, and why Rippit reads 100% of conversations. Brex’s public story describes rebuilding its QA playbook from the ground up.

Comparing QA tools? See how Rippit stacks up against Zendesk QA, Level AI and Solidroad.

FAQ: QA for 100% of support conversations

Does full-coverage QA mean every score is automatically correct?

No. Full coverage solves the coverage problem; calibration solves measurement; humans handle accountability. A bad rubric applied to 100% of conversations is still a bad rubric.

Can AI QA scores be used for performance reviews?

Use caution. Before an AI-derived criterion feeds a formal employment decision, validate it against experienced reviewers and set up a process for disputed scores.

Do you still need manual QA?

Yes, but much less random manual scoring. Humans stay on calibration, coaching, appeals and improving the rubric.

Is this different from what MaestroQA did?

MaestroQA is Rippit’s former name, not a different product; QA customers use the same platform. QA is now one of several agents teams build on the same data.

How long does it take to get the first scores?

Connect your customer conversation hubs to Rippit in one click—including Zendesk, Intercom, Gong and more—then describe the agent in plain language; it scores on a schedule. Through MCP, the agent can also read, write and update data across your broader stack, including Salesforce, HubSpot, Slack, Notion, Guru, Jira, Snowflake, and more. Budget more time for calibration than setup.

Where conversations become

insights

actionable data

business intelligence

enterprise visibility

insights

I fear not the man who has practiced 10,000 kicks once, but I fear the man who has practiced one kick 10,000 times
Peloton
legal zoom