Playbook

AI-Powered Conversation Taxonomies Playbook

How to turn every customer conversation into structured insight

All conversations tagged, like magic

Download the ebook to share with your team.

Your team is sitting on tens of thousands of conversations. Intercom tickets, Gong calls, chat logs.

And nobody can answer the simplest question about them: What are customers actually telling us?

The old answer was a tagging taxonomy: a tree of categories, applied by agents, one tag at a time. The new answer looks the same on the surface and works completely differently underneath. This guide covers what a conversation taxonomy is, why AI didn't just automate the old playbook but inverted it, and how to build one across 100% of your conversations.

Everything in the customer voice and FAQ sections below comes from a full year of real conversations with CX, quality, and insights leaders, verified against the source transcripts.

Key Takeaways

  • A conversation taxonomy is a structured classification system for customer conversations (themes, intents, root causes, sentiment) applied consistently at scale.
  • AI builds taxonomies bottom-up from real conversations instead of top-down from a conference-room guess.
  • AI classification is multi-dimensional: one conversation, ten-plus dimensions, no extra agent effort.
  • The same taxonomy should span support tickets and sales calls. It's all the voice of the customer.

What a conversation taxonomy is

What is a conversation taxonomy?

A conversation taxonomy is a structured classification system that organizes customer conversations (support tickets, sales calls, chats) into consistent categories such as themes, intents, root causes, and sentiment, so teams can measure, trend, and act on what customers are saying at scale.

Tags and taxonomies are not the same thing. A tag is a label; a taxonomy is the organized system of labels: the hierarchy, the definitions, the rules for when each one applies. A pile of 121 helpdesk tags that includes "tag", "tag1", "no", and "" (we've seen it) is not a taxonomy. A three-level structure is:

Level 1 (Theme)
Level 2 (Sub-theme)
Level 3 (Root cause)
Billing
Failed payment
Expired card
Billing
Failed payment
Bank declined
Billing
Unexpected charge
Plan auto-upgrade

What is an AI-powered conversation taxonomy?

An AI-powered conversation taxonomy uses large language models to read full conversation transcripts and do two jobs at once: discover the category structure from the conversations themselves, and apply that structure to every new conversation automatically, across multiple dimensions (intent, root cause, sentiment, churn risk) rather than a single tag.

Why conversation taxonomies matter

Unstructured conversation data is where insight goes to die. Surveys capture under 10% of customers. Manual QA samples 1–5% of interactions. The rest, the overwhelming majority of what customers say, sits unread, while roadmaps get driven by whoever escalated loudest.

Customers describe the gap in almost identical numbers:

Right now, we're only able to manually QA about two percent of tickets… we can assume that if something comes up in QA that it's probably a bigger pattern, but we don't have the data to substantiate that.

support leader, healthcare-software company

It was over $20,000, not even 2% of my cases were QA'd… it was statistically irrelevant data.

customer-service leader, online book retailer

We evaluate the conversation in a very old-fashioned way… block out the conversation ID, put it in a separate Excel sheet where we store the data on what kind of chats we've checked so far: a few percent of tens of thousands of monthly interactions.

head of performance, multi-brand iGaming operator

The cost of doing this by hand

The cost of doing this by hand is real. Conversation analytics used to be a professional-services project: consultants built custom taxonomies and bespoke models per company, which is why legacy enterprise tools reportedly carried minimum contracts around $250,000.

$250,000

reported minimum contract

And agent-entered disposition codes are notoriously inaccurate anyway. At the top of the market the pain compounds: a business strategy lead at a consumer financial-services company with roughly two million associate-handled calls a month put it plainly:

It's so manual today, and because it is lagging it just takes forever to figure out what the issue might be and then go root cause it.

Who benefits when the taxonomy actually works

Who benefits when the taxonomy actually works: CX and support leaders get contact-driver reporting, deflection targets, and QA coverage on every ticket instead of a sample. Sales leaders get objection, competitor-mention, and lost-deal trends straight from calls. Product teams get feature-request volume, bug themes, and churn root causes quantified instead of anecdotal.

Manual Tagging Is Broken

Manual tagging is broken. AI doesn't just fix it, it inverts it.

Manual tagging fails predictably. Agents skip tags under time pressure. Two agents label the same issue differently. One tag per ticket flattens a nuanced conversation into a single word. And the taxonomy itself goes stale the day the product changes.

Keyword automation isn't the fix either. Keyword rules flag "cancel" in "how do I cancel a duplicate order" as churn.

In one real evaluation, keyword-based churn flagging returned ~1,100 calls when the true number was 15–20.

Keywords match words; taxonomies need judgment.

From top-down guessing to bottom-up discovery

Manual taxonomies are designed in a conference room, then reality gets forced into them. One retailer told us their product insights "reflected internal workflows more than the customer's experience." AI reverses the flow: it reads the actual corpus, surfaces the categories that exist in it, and humans curate and name them.

From one tag to many dimensions

An agent closing a ticket picks one tag. An LLM reading the transcript can classify intent, theme, root cause, sentiment, urgency, churn signal, and resolution status in a single pass. A custom question becomes a custom dimension, asked once and answered across 100% of the corpus.

This also retires the classic advice to "keep your taxonomy under 30–50 tags." That rule was a human working-memory constraint, not an insight constraint. When no human has to memorize the tag list, depth and precision stop being the enemy.

One insights manager at a major news media company runs a tiered taxonomy of roughly 160 granular issue tags laddering up to product-level tags, across about a million interactions a year on six channels:

We developed over the years… a tiered taxonomy… about 160-something granular issue tags.

Manual tagging vs. AI-powered taxonomy

Customer Voice

Real-world taxonomy use cases

Every use case in this section comes from a real conversation with a CX, quality, or insights leader over the past year, describing what they want to do, or are already doing, with their conversation data. Quotes are verbatim (lightly cleaned for filler words), verified against the source transcripts, and anonymized by role and industry.

Contact drivers you can actually drill into

Helpdesk categories tell you that volume went up, not why.

I can recall all the tickets that were related to that new product, and then I can sort them by the contact reason. But from there, that's where it ebbs. You can't get to the granular.

director of customer care, consumer pet-tech COmpany

If the intent was to schedule an appointment and you could not, then why not? … These four reasons make up the 100 percent of why not. So we don't have that today.

patient-access leader, academic health system (about 300 people handling well over 100,000 calls a month)

Churn-reason classification by product line

We’re building reporting on what are the specific churn reasons for customers that are specifically using [one product], and are they product limitation or service based.

CX quality lead, HR software company

There's 80 percent of tickets on average that do not receive a survey, either because we may not have one attached to that particular issue type, or players decide not to give us a survey, or it times out. We're finding that there could be value in those 80 percent.

Quality manager, mobile gaming

Predicting churn better than your agents can

A buy-now-pay-later fintech turned conversations into a churn model, and benchmarked it against humans:

We have a model that predicts 80 percent whether they're going to churn or not based on the conversation. We also did a contest with our agents (hey, predict whether they're going to churn based on your conversation) and they had like a 30 percent success rate.

buy-now-pay-later fintech

A retention leader at a fiber ISP uses classification to level-set the save opportunity:

Out of this pool (cancellation request) we saved four. And of that pool we attempted to save 104… as just a level set of what we're looking at.

retention leader, fiber ISP

A CX quality manager at a robo-advisor used classification to answer a board-level question, typing this request into their workspace:

I need insights into why customers chose to defund and/or close their accounts… and what percentage of that churn can be reasonably attributed to the [recent] security incident.

CX quality manager, robo-advisor

Indirect churn signals and save plays

I want to pull the tickets that maybe had two dings for a direct churn risk or even indirect churn risk, build a report off of that, have our support team follow up, and then three months down the road show the executives that these members are still enrolled because we followed up.

voice-of-customer analyst, virtual healthcare provider

Root causes ranked by cost to serve

I would love to better split our data by root causes… to find the exact customer problems that take us the longest to solve, and then see if we can make a business case to build better solutions into the product.

service quality lead, HR software company

A CX leader at a sports-betting operator did exactly that for promotions:

[One promo type] is growing as a share of volume and customers are getting more pissed off about it… so we pulled in 10,000… it broke down exactly what about every promo went wrong.

CX leader,sports-betting operator

Failure-mode triage: people, process, or product

We want to understand the whole customer experience and all of the points where we're failing, and then break them down into which areas we need to fix: is it the training gaps, is it the knowledge base, is it just our process in general?

CX lead, online gaming company

Product feedback with revenue attached, and decisions gated on counts

A CS operations lead at a talent-acquisition SaaS company classified a year of cases against a single product gap:

I was basically able to extract that we had 288 cases… and turn this into a $65 million ARR impact point for 138 customers: find the top five themes, provide that back to the product manager directly.

CS operations lead, talent-acquisition SaaS company

The same logic runs in reverse, using counts to unblock decisions.

We just want to implement a change, but without knowing how many calls are coming in for this certain call type, we cannot make a decision.

collections QA lead, BNPL company

An edtech team even stood up a classifier before the event it measures:

We are actually removing [a] feature tomorrow from our website and we want to know how many customers reach out to us about this feature removal.

Edtech Team

Emerging-issue detection before agents flag it

We're thinking about better incident detection, ways that we can notice trends as they're happening. Here's all the ones we've tagged as known incidents. Can you find ones that don't match this criteria?

support operations admin, fintech banking platform

A CX operations analyst at a healthcare platform explained why speed matters:

By the time it's flagged by an agent, five, six, seven days have passed already, and we knew that we could find these earlier based off the new tickets coming in.

CX operations analyst, healthcare platform

Her team now runs classifiers on open tickets to:

Identify emerging trends prior to a human agent intervention… to try and jump ahead of the dissatisfaction.

CX operations analyst, healthcare platform

Sometimes the outside world moves first. A ticket-resale marketplace discovered a resale-policy change only when customers wrote in:

We're not finding out about it until fans are contacting our support team.

ticket-resale marketplace

They then used classification over tens of thousands of tickets to size the impact.

A global CX insights manager at an anime-streaming service used AI translation plus classifiers to monitor a brand-new Thai-language launch:

The team is trying to fix any technical issues in as real time as possible based on what customers are saying. Our help desk doesn't translate Thai well…

global CX insights manager, anime-streaming service

Emerging-issue detection before agents flag it

This is where sampling isn't just inefficient. It's exposure.

The goal is to build an LLM that can identify whether a customer is vulnerable according to FCA regulations and return a simple 'yes' or 'no'

software engineer, retail-investing platform

A CX manager at a sports-betting operator needs context, not keywords, for responsible gaming:

If a customer says, 'I can't pay my rent,' that is a Responsible Gaming concern, whereas if someone says 'I play within my financial means and my rent is on time,' that is not a concern.

CX manager, sports-betting operator

A senior quality manager at a pharmaceutical company runs classification against a regulatory clock:

Adverse events need to be reported within one business day… if it took 24 hours we would be out of compliance.

senior quality manager, pharmaceutical company

A quality assessor at a travel-benefits company caught an agent taking card details on a recorded line during manual QA, and asked the obvious next question:

Is there something we can build for the AI to look out for that across all calls automatically?

Quality assessor, travel-benefits company

And a program manager at a crypto exchange flipped retention policy into a classification problem: identify the recordings where customers mention equities trades, which must be kept seven years, instead of retaining everything.

From sampling to full QA coverage, and QA that becomes coaching

Once classification reads every conversation, the sampling ritual starts to look strange. One QA leader asked it outright:

Why are we having a human measure that when the AI is already doing it across every single interaction?

QA Leader

The manual hours don't disappear. They move up the value chain, from grading logistics to coaching interventions.

Sentiment that understands nuance

Generic sentiment isn't a taxonomy. A productivity-software team caught the difference:

The support bot says 'Did I resolve your issue?' and the user answers 'no.' The classifier is saying this is 'negative' sentiment, when it is factual, not emotive.

productivity-software team

A QC lead at a supplements retailer built the opposite of a complaint detector, a turnaround detector:

The purpose of this LLM is to catch tickets with high customer dissatisfaction that got positive CSAT… we are checking if agents can be somehow rewarded based on this.

QC lead, supplements retailer

And a clinical quality lead at a longevity health clinic found their most emotionally intense churn driver hiding in visit transcripts:

Come to find out, biological age and lab results represent the highest intensity of member frustration.

Clinical quality lead, longevity health clinic

QA-ing your own AI agents

The newest taxonomy use case isn't about humans at all. Teams now classify their chatbots' conversations to audit them, including auditing other AI.

A product support QA specialist at a design-software company uses conversation analysis to check an AI model that auto-applies their ticket taxonomy:

The team is iterating on the taxonomy weekly based on audit findings…

product support QA specialist, design-software company

Bot-vs-human attribution, escalation detection, and "did the bot actually resolve it" are becoming standard taxonomy dimensions.

Where does this end up? One sports-betting CX leader has been circling the logical conclusion for months:

Is there a future state where we completely get rid of category and subcategory and just have issues: trend existing issues, recognize new issues? That seems ripe for something like this.

CX leader, Sports-betting platform
The pattern across all twelve: nobody wants tags for their own sake.

They want the taxonomy as the bridge from "we have conversations" to "we made a decision."

How to build it

How to build an AI-powered conversation taxonomy (step by step)

Common pitfalls

Over-granular third levels

Splitting until every branch holds a handful of conversations, so no level has enough volume to report on.

No catch-all category

Ambiguous conversations get force-fit into a real category, and every count quietly drifts away from what customers said.

Freezing it after launch

The product keeps moving and the tag list does not, so new issues land in categories written for last year.

One more, from the field: don't stop at the most common themes. The exceptions and outliers inside a theme are often where the future shows up first.

FAQ

These aren't hypothetical questions. They're the ones CX and quality leaders actually asked most often across a year of real conversations, ranked roughly by how frequently they came up.

How many categories should a conversation taxonomy have?

Three to five levels deep, with as many categories as your decisions require. The old "30–50 tags max" rule existed because humans had to memorize the list; AI classification removes that constraint. The real limit: every category should have an owner and a use.

What's the difference between a tag and a taxonomy?

A tag is a single label applied to a conversation. A taxonomy is the structured system of labels (hierarchy, definitions, and rules) that makes tags consistent, comparable, and reportable across teams and time.

Can AI automatically tag support tickets?

Yes. LLM-based classifiers read the full transcript of every ticket and apply your taxonomy automatically, replacing agent-entered dispositions, which are slow, inconsistent, and expensive, with consistent classification across 100% of conversations.

Why does AI give different answers when I ask the same question twice?

Open-ended AI queries are non-deterministic. As one analytics user put it: "I ran queries two times. It gave me two different results. Why?" The fix is architectural: classify per-row with fixed category lists for anything you'll report on repeatedly, and reserve free-form AI chat for exploration.

How accurate is AI classification, and how do I validate it?

Validate against human judgment on 25–100 hand-reviewed conversations per category until alignment stabilizes, including deliberately hard cases, not just random samples. Teams typically require higher alignment for compliance classifiers (90%+) than for subjective dimensions like sentiment (~70%), and keep an "unknown / not enough information" bucket rather than forcing a guess.

How do you stop the AI from inventing categories?

Constrain the output to an exact list of labels with a catch-all option, rather than free text. Unconstrained columns drift into near-duplicate categories. Then spot-check supporting evidence: a classifier should be able to point at the passage that justified its label.

Can an agent dispute an AI-generated QA score?

They should be able to. This is the most common trust question teams ask, especially where scores feed KPIs or bonuses. Best practice is a human appeal loop: agents flag disagreements, reviewers grade the disputed examples, and that feedback refines the classifier's instructions.

Can AI classification run retroactively on historical conversations?

Yes. That's one of its biggest advantages over manual tagging. A new category can be applied to months or years of history in one enrichment run, so you get a trendline on day one instead of waiting for data to accumulate.

What happens to my taxonomy when the underlying AI model changes?

This is the quiet governance problem: one team saw a company-level sentiment metric drop sharply after a silent model migration and spent weeks investigating. Version your taxonomy against model versions, validate new models side-by-side on the same conversations before switching, and annotate reporting when either changes.

Does AI classification work across languages?

Yes. Modern LLMs classify conversations in most major languages from English-language instructions. One team monitored a brand-new Thai-language launch this way when their help desk couldn't even translate the tickets. Watch for translation-layer artifacts (e.g., agent and customer appearing to speak different languages) and note them in classifier instructions.

How is this different from Intercom Topics or Gong Smart Trackers?

Built-in tools classify one platform's conversations with largely preset, single-dimension categories. A true conversation taxonomy is custom to your business, multi-dimensional, and spans every source (support tickets, sales calls, and chat) in one consistent structure.

How often should you update a conversation taxonomy?

Treat it as a living system: curate quarterly, and rely on AI to continuously surface emerging themes that don't fit existing categories. Version categories when definitions change so historical reporting stays comparable.

Do I need training data or ML engineers?

No. Modern LLM classification is instruction-based, not model-training: you describe the category in plain language, test it on real conversations, and refine the wording. Classification works out of the box and is refined with a prompt.

Conclusion

The taxonomy is no longer imposed. It's discovered

For decades, a taxonomy was something you imposed on conversations and hoped agents complied. AI inverted that: the taxonomy is now something you discover in your conversations, and apply to all of them, every dimension at once, at 100% coverage.

Rippit connects to Intercom, Gong, and 25+ other conversation sources, discovers your taxonomy from real conversations, and runs AI Classifiers on every new one: live KPIs on churn, sentiment, contact drivers, and anything else you can describe in a sentence. Teams like Checkr, Brex, and Klaviyo use it to go from sampling a sliver of conversations to reading all of them.

See your first taxonomy built on your own data.  Request a demo →

Table of Contents

Get the ebook

Share this post

Stop asking, start reading. Uncover true customer sentiment.

Replace biased 5% survey samples with a satisfaction score on every conversation.

All things are one. When we perceive this, we see that the flowers, the trees, and the stars are all part of our own body."

Where conversations become

insights

actionable data

business intelligence

enterprise visibility

insights

I fear not the man who has practiced 10,000 kicks once, but I fear the man who has practiced one kick 10,000 times
Peloton
legal zoom