At 50,000 conversations a month, the answer to “why are customers contacting us?” usually comes from a helpdesk dropdown picked under handle-time pressure.
The data isn’t missing. It’s just not a measurement of anything.
Key takeaways
- Agent-applied tags fail at volume in four predictable ways: taxonomy drift, category collapse, first-touch bias, and an “Other” bucket that absorbs everything new.
- Before acting on any driver list, run two tests: hand-label 50 conversations and check where people disagree, then re-run the classification and check the numbers hold.
- A driver list is only useful with trend and example conversations attached.
Why do support tags stop working at scale?
1. Taxonomy drift
Two agents read the same conversation and pick different categories. A payment that failed because a card expired is Billing to one agent and Account to another. Both are defensible.
2. Category collapse
One broad tag becomes the default. “Login issue” absorbs password resets, SSO failures, locked accounts. Your largest category becomes your least informative one.
3. First-touch bias
The tag reflects what the customer said the problem was. A conversation that opens with “Where is my order?” and ends with a broken tracking page gets counted as Order status: the symptom, not the root cause.
4. The “Other” sink
Anything genuinely new has no tag yet, so it lands in Other or the closest bad fit. Your newest problem is structurally invisible.
What should you measure instead of support tags?
Classify the actual conversation, on several fields at once, across the full dataset:
- Why the customer contacted you, and what they were trying to accomplish
- The underlying root cause, and the product or workflow involved
- Whether it was resolved, repeated or preventable
- How sentiment changed
One payment conversation might carry contact driver Payment failure, root cause Expired card, product area Mobile checkout, repeat contact Yes and resolution Unresolved. Reducing that to Billing throws away most of what you learned.
Can you change the taxonomy later?
Yes, if you keep the conversations, not just the labels.
Split Billing → Payment problem into expired card, bank decline and processor error, and old tagged tickets don’t follow. Someone retags them, or history breaks.
When the original conversations stay in a dataset, you can change the definition and reclassify historical conversations using the new logic. Your taxonomy can evolve with the business without losing the past.
What should a contact-driver report contain?
For each driver in a ranked list, include:
| Include | Why it matters |
|---|---|
| Volume and share | How large is the issue? |
| Change vs. prior period | Is it getting better or worse? |
| First-seen / emergence | Is this an established problem or something new? |
| Resolution/outcome | What happens to these customers? |
| Example conversations | Can someone verify what the category actually means? |
| Driver | Volume | Share | Vs. prior |
|---|---|---|---|
| Payment failure · expired card | 4,210 | 8.4% | +12% |
| Login · SSO failure First seen this month | 3,180 | 6.4% | New |
| Order status · broken tracking page | 2,940 | 5.9% | −5% |
| Refund timing | 2,120 | 4.2% | +1% |
If nobody can open the conversations behind a number, the meeting becomes an argument about the metric.
Traceability turns an AI-generated insight into something a business can trust.
How do you validate AI-generated contact drivers?
Run two tests before you act on any driver list, including one an AI produced.
Test 1: Human disagreement
Have two experienced people independently classify 50 conversations against the same definitions. If disagreements cluster in some categories, fix those definitions before automating.
Test 2: Classification stability
Rerun your classifier against a known validation set. If classifications move without an intentional change to the definition or model, investigate. The goal isn’t 100% determinism. It’s enough reproducibility that a movement in your dashboard reflects customer behavior, not classification behavior.
Why analyze all conversations instead of a sample?
Sampling is useful for validation.
Your largest contact driver will probably appear in a sample. A new problem, or a low-volume issue that drives churn or compliance risk, may not.
How does Rippit analyze contact drivers?
Rippit treats your customer conversations as a dataset.
Connect your customer conversation hubs to Rippit in one click—including Zendesk, Intercom, Gong and more. Through MCP, the agent can also read, write and update data across your broader stack, including Salesforce, HubSpot, Slack, Notion, Guru, Jira, Snowflake, and more. Rippit handles ingestion, then enriches each conversation with structured and AI-derived fields. A non-technical user defines the analysis in plain language. The resulting fields become reusable parts of the dataset, so next month’s question doesn’t start from scratch.
Who is this approach best for?
- CX, support ops and product teams at 10k–100k conversations a month whose tagging taxonomy has drifted
- Leaders who need a defensible answer to “why are customers contacting us?”
- Teams whose conversations live in Intercom, Zendesk or the other tools Rippit connects
- Teams under a few hundred conversations a month, where reading them is faster and more accurate
- Teams that only run surveys and have no conversation volume to analyze
- Teams that want a general BI tool rather than analysis of conversation text
Comparing customer insights tools? See how Rippit stacks up against Enterpret, Unwrap and Zendesk.
What is Rippit?
Rippit is a conversation data platform that lets anyone build AI agents using all of their conversation data. You don’t have to be an engineer. Connect your customer conversation hubs to Rippit in one click—including Zendesk, Intercom, Gong and more—describe what your agent should do in plain language, and it reads 100% of your conversations and delivers results on a schedule. Through MCP, the agent can also read, write and update data across your broader stack, including Salesforce, HubSpot, Slack, Notion, Guru, Jira, Snowflake, and more.
FAQ: finding out why customers contact support
Can’t we just fix our tagging taxonomy instead?
Improve it, but tags stay a quick judgment by a rotating set of people, so drift and first-touch bias return. Classifying the conversation text removes the data-entry step and lets you re-derive drivers historically when definitions change.
How many conversations do we need to analyze?
All of them. A sample gets your biggest driver roughly right. Rare, severe and brand-new issues live in the long tail, and a sample is designed to exclude it.
How long does it take to get a first driver list?
Connect your conversation source in one click, describe what you want the agent to find in plain language, and it reads your conversations and delivers results on a schedule.
Does this replace CSAT or NPS surveys?
No. Surveys hear from the customers who respond; conversation analysis covers everyone who contacted you. Most teams run both.
Can we ask these questions from ChatGPT or Claude?
Yes. Rippit hosts its MCP server, so an admin adds one connector URL and there is nothing to deploy. Each person sees only what they can already see in Rippit, and the assistant reasons over fields already derived from every conversation.
What’s the bottom line?
The conversation is the source data. Keep it, structure it, and make every number traceable to the conversations behind it. Then you’re answering from what customers actually said, not from a dropdown.
.webp)