The reliable way to use AI on customer feedback is classification, not summarization. Every review, ticket, call excerpt and survey answer gets tagged against a fixed set of themes, with the exact customer quote attached. That gives you counts you can trend, language you can use in copy and a result you can actually check.
This expands the one-paragraph feedback-mining use case from AI agents for marketing into a full pipeline, from collection to the monthly handoff.
What voice-of-customer data you already have
Most companies don’t lack feedback. They have plenty, spread across five tools owned by four teams, and nobody reads all of it.
| Source | Typical home | Best for | Watch out for |
|---|---|---|---|
| Public reviews | G2, Capterra, Trustpilot, Google, store review apps, Amazon | How buyers describe value and compare you | Skews to extremes; incentivized reviews |
| Support tickets and chat | Zendesk, Intercom, Gorgias, HubSpot | Friction and confusion after purchase | Agent replies mixed into the text |
| Sales-call transcripts | Call recorders and meeting-note tools | Pains, alternatives and objections before purchase | The rep does most of the talking |
| Surveys and NPS | Post-purchase, in-app and onboarding surveys | Reasons behind scores | Short answers, leading questions |
| Cancellation reasons | Exit surveys, billing tools, CS notes | Why customers leave | Dropdown picks hide the real reason |
| CRM notes | Deal and account records | Lost-deal context | Paraphrased by reps, not customer words |
Each source covers a different moment: calls and competitor reviews cover the decision, post-purchase surveys cover the purchase, tickets cover use and exit surveys cover the end. B2B teams usually get the most from call transcripts and churn notes; DTC brands from product reviews, post-purchase surveys and support tickets. Start with the two or three sources with the most volume and add the rest once the pipeline runs.
Collecting and cleaning sources
The goal is one table where every row is a single piece of customer feedback. A scheduled workflow (n8n, a small script or a warehouse job) pulls from each source’s export or API into it.
- Use a fixed schema. Comment ID, source, date, text, rating or score if any, a link back to the original, and segment fields: plan or product, customer tenure, deal stage or won/lost, acquisition channel.
- Split long documents into units. A 45-minute call is not one comment. Keep only customer turns, using speaker labels, and break them into topic-level excerpts. For tickets, keep the customer’s messages and drop agent replies and auto-responses.
- Strip personal data before any model call. Replace names, emails, phone numbers, addresses and order numbers with placeholders.
- Remove noise. Duplicates, one-word replies, spam reviews and internal test entries.
The segment fields are what make the analysis useful later. “Setup effort” appearing in lost deals from your best-fit segment matters more than the same theme from customers who were never a fit.
Building a theme taxonomy
A taxonomy is the fixed list of themes every comment gets sorted into. Build it by reading before you prompt. Hand-read a sample of 150-300 comments across sources and note themes as they appear. You can have a model propose candidate themes from a sample to speed this up, but a person who knows the customer edits and owns the final list.
Organize themes under categories that map to decisions:
- Pains and jobs: why they went looking for a solution
- Outcomes: the value they describe getting
- Objections: what almost stopped them buying
- Friction: product, UX, delivery or service problems
- Alternatives: what they used or considered instead
- Requests: features, sizes, integrations they ask for
Give each theme a name, a one-sentence definition, include and exclude rules, and two example quotes. Keep the whole list to roughly 25-50 themes, and never create themes that are only sentiment, such as “positive feedback.” Sentiment is a separate field.
| Category | Theme | Definition | Illustrative quote |
|---|---|---|---|
| Pain | Manual reporting | Time lost compiling reports by hand | “Every Monday goes to pulling numbers into a sheet” |
| Objection | Setup effort | Worry about implementation time or needing IT | “We don’t have anyone to set this up” |
| Friction | Delivery delay | Order arrived later than promised | “Took two weeks longer than the site said” |
| Alternative | Spreadsheets | Current workaround is spreadsheets | “Right now it all lives in Excel” |
The quotes above are made up to show the format.
Prompting for classification, not summaries
Summaries fail as analysis. They compress to whatever sounds most prominent, lose the counts, can’t be compared month to month, and often produce paraphrases that read like quotes. Classification turns each comment into structured data you can count outside the model.
A classification prompt has four parts:
- The taxonomy, with theme IDs, definitions, include and exclude rules
- The comment, with its ID, source and segment fields
- Rules: assign zero to three themes from the list only; if nothing fits, return “other” with a suggested theme name; copy the quote exactly; don’t infer beyond the text
- A strict output format, such as JSON
{
"comment_id": "tkt-4821",
"themes": ["objection.setup_effort", "alternative.spreadsheets"],
"sentiment": "negative",
"quote": "we don't have anyone to set this up",
"suggested_theme": null
}
Then validate in code, not by eye. Reject any theme ID not in the taxonomy, and check that the quote appears word for word in the original text after normalizing case and whitespace. That substring check catches invented and paraphrased quotes; it won’t catch a real quote under the wrong theme, which is what the accuracy check below is for. Send one comment or a small batch per call, keep temperature low where the model allows it, and pin the model version so month-to-month changes come from customers, not the model.
Counting happens in a spreadsheet or database: themes by source, by segment, by month. Summaries still have a place, at the end: once the counts exist, a model can draft the monthly narrative from the counts and verified quotes you hand it, and nothing else.
Checking accuracy against a sample
Before trusting the counts, hand-code 100-200 comments, spread across sources, and run the model on the same set. Have a second person code part of the sample too. If two people disagree on a theme, the definition is the problem, not the model.
Compare results theme by theme, not as one overall score. A model can be right on most comments overall and still badly miss the one theme you care about.
| What you see | Likely cause | Fix |
|---|---|---|
| Model over-tags a theme | Definition too broad | Add exclude rules and counterexamples |
| Model misses a theme | Customers phrase it in ways your examples don’t cover | Add examples in customer language |
| Two humans disagree | Theme is ambiguous | Split, merge or redefine it |
| Large “other” bucket | Gaps in the taxonomy | Review “other” and add themes |
| Quotes fail the substring check | Model paraphrasing | Tighten the rule; drop failures from reporting |
Act on themes that clear your accuracy bar and label the rest directional. Re-run the full check whenever the taxonomy, prompt or model changes.
Turning themes into messaging, CRO and product input
The monthly output is a short report: top themes by volume and trend, broken out by segment, each with three to five verbatim quotes, plus anything new or rising. Weight themes by who says them. A theme from lost deals in your best segment outranks a louder one from low-fit customers.
Then route each theme to the team that can act on it:
- Messaging. Pains and outcomes, in customer words, become headlines, ad angles and email subject lines. They’re also the evidence for testing your positioning; see how to write a positioning statement your sales team will use.
- CRO. Objections become page content and test hypotheses. If “setup effort” leads the sales-call objections, add an implementation timeline and a support promise near the demo button and test it. Delivery or sizing friction on a DTC store goes onto the product page.
- Product and ops. Friction and requests go to the backlog with counts, segments and quotes, which carry far more weight than “customers keep asking for.”
- Sales and CS. The objections that recur, with the phrases prospects actually use, become talk tracks.
Every theme in the report gets an owner and either one action or an explicit “no action this month.” Without that, the report becomes something people skim.
Keeping the pipeline running monthly
Once built, the pipeline typically needs a few hours of human time a month. Most of it goes to review, not operation.
- New data pulled from every source on schedule
- Personal data stripped before any model call
- Classification run on the current taxonomy version, with the version logged
- Quote substring check passed; failures reviewed
- Spot-check of 20-30 comments against human coding
- “Other” bucket reviewed; themes added or merged
- Report sent with owners and actions
- Last month’s actions reviewed: did the related theme move?
When you change the taxonomy, re-run the last few months on the new version so trends stay comparable. This kind of pipeline is a standard build in my AI automation work: the collection workflows, the classification step with validation, and the monthly report your team actually reads.
Get it built
If your reviews, tickets and call recordings sit in separate tools and nobody reads them together, I can build the pipeline and run the first monthly cycles with your team. The AI automation add-on starts at $2,500/mo, or begin with a Growth Audit, $1,500 fixed and credited if we continue. See pricing or get in touch.