The pile you never quite get through
Every product manager has one. Support tickets tagged "feature request". Notes from sales calls. Free text from a satisfaction survey. A spreadsheet of things customers asked for at the last user group. Churn reasons typed in a hurry by account managers. Somewhere in that pile are the three problems that should shape your next quarter, and you have read perhaps a fifth of it.
A language model can read all of it in minutes, group it, name the groups and write a summary your leadership team will find persuasive. That is a real gain. It is also the exact point where the tool is most likely to put something on your roadmap that no customer ever said.
This lesson is about getting the first without the second. If your material is mostly long-form interviews rather than tickets and feedback, AI for UX and Product Design has a lesson devoted to interview synthesis and verbatim quote checking. Here the focus is the messier, higher volume material product managers live with, and the decisions it feeds.
Four ways feedback synthesis goes wrong
Invented evidence. A theme with no real support, or a quote that reads beautifully and appears nowhere in your data. It fits the pattern of a finding, so the model produces it.
Invented counts. "Around a third of customers mention onboarding" sounds like analysis. Unless you asked it to count and then checked, it is a guess shaped like a number. Models are unreliable at counting across long material, and they will not tell you they estimated.
The request mistaken for the problem. Customers ask for solutions: an export button, a dark mode, an integration. A naive synthesis groups by what was asked for. What you need is grouping by the underlying problem, because five different requests can come from one frustration.
Tidiness. Summarising pulls towards a clean story. The customer who said the opposite, the segment that does not fit, the complaint that only enterprise accounts raise: all of it gets smoothed away unless you ask for it by name.
The loudest feedback is not the most common feedback. A single customer who files twenty tickets can become a "theme" if you do not ask the model to count distinct customers rather than mentions.
Make every theme traceable
The discipline is the same one that governs any research: no theme survives unless it points at specific items you can go and look at. Give each piece of feedback an identifier before you paste it in. Ticket numbers, row numbers, a short code per call note. Then insist that every claim cites them.
❌ Weak prompt
Prompt
Here is our feedback from last quarter. What are customers asking for and what should we build?
Output
Customers are asking for better reporting, more integrations and a simpler onboarding experience. Reporting is the most requested feature, with many users saying things like 'I need to see my data at a glance.' We recommend prioritising a new reporting dashboard.
No identifiers, no counts, a quote you cannot find, and a jump straight from requests to a build recommendation. 'Most requested' is unverifiable, and the dashboard suggestion is the generic answer for almost any product. You would have to reread everything to know which parts are real.
✅ Good prompt
Prompt
Below are 140 feedback items, each starting with an ID such as T-0012 or CALL-07, followed by the customer account and plan. Group them by the underlying problem the customer has, not the feature they asked for. For each theme: list every item ID in it; count distinct customer accounts, not items; quote two items word for word with their IDs; and add a COUNTER line listing any items that contradict the theme or show customers who are happy with the current behaviour. Put anything that fits no theme under UNGROUPED. Do not recommend what to build. Do not estimate: if you are unsure whether an item belongs, list it under UNSURE.
Output
Theme A: cannot tell whether the weekly import worked. Items: T-0012, T-0031, T-0044, CALL-03, CALL-11, T-0102. Distinct accounts: 5 (T-0031 and T-0044 are the same account). Quotes: T-0012 'no idea if it ran or not, had to check by hand'; CALL-11 'we found out a week later it had failed'. Asked-for solutions in this theme vary: email alerts, a status page, an export log. COUNTER: T-0090 says the import 'just works, never think about it'. UNSURE: T-0077 mentions imports but may be about speed.
Every claim now points at something you can open. Grouping by problem shows that three different requests share one cause. Counting accounts rather than items stops one noisy customer becoming a trend, and the UNSURE list gives the model a legitimate place to put doubt instead of forcing a guess.
Then check it. Open a sample of the cited items, including every quote. Recount the accounts for the two or three themes that might change a decision. This takes perhaps twenty minutes, and it is the difference between synthesis and fiction.
Checkpoint
Label every feedback item before pasting it, and accept no theme, quote or count that does not cite items you can open and check yourself.
Volume, weighting and what the pile cannot tell you
When the pile is too big. If you have more feedback than fits in one conversation, do not let the model summarise in batches and then summarise the summaries. Each layer drifts further from the source. Instead, synthesise one batch with IDs, then give the next batch alongside the existing theme list and ask it to assign items to existing themes or propose new ones, still citing IDs. Context Engineering explains why long sessions drift and how to structure work around it.
Weighting is yours. Five small accounts and one large account with the same problem are not the same fact. You can include plan or segment in each item so the output can be sliced, but deciding what a segment is worth is a business judgement the model has no basis for.
Absence is invisible. Feedback comes from customers who stayed long enough to complain. The people who churned quietly, or never signed up, are not in the pile. A synthesis can tell you what your vocal customers want. It cannot tell you what your market wants.
Before sharing a synthesis, ask one more question in the same conversation: "Which of these themes could be explained by who gives us feedback rather than by what customers experience?" It is a cheap way to catch the bias built into the source.
Checkpoint
Group by the underlying problem rather than the requested feature, count distinct customers rather than mentions, and remember the pile only contains people who chose to speak.
📝 Quiz
Question 1 of 4A synthesis says 'around a third of customers mention onboarding'. You never asked it to count. What should you assume?