LearnAI home

Research and Synthesis · Lesson 2

Synthesising Research Without Inventing It

Turning interviews and feedback into themes, and catching the findings it made up.

The stage where it earns its keep, and the stage where it lies

Synthesis is the part of research that takes the longest and looks the least like work. Twelve interviews, a few hundred support tickets, a survey with a free-text field nobody wanted to read. Somewhere in there are four or five things worth telling the team. Getting from one to the other has always been slow.

A language model is genuinely good at this. Give it a pile of transcripts and it will cluster them, name the clusters sensibly, and write the whole thing up in the register your stakeholders expect. What would have been two days of index cards becomes an afternoon.

It is also, at exactly this moment, at its most dangerous. Here is the thing you must internalise before you use it on real research:

It will produce a theme that no participant expressed, and a quote that nobody said, in the same confident tone as the themes and quotes that are real.

Not occasionally. Not as a glitch. It is a direct consequence of what the tool is: something that produces text fitting the pattern of research output. A fabricated quote fits that pattern beautifully. It will be well phrased, emotionally plausible, and attributed to P7.

Why fabricated findings are so hard to catch

Ordinarily you catch errors because they look wrong. Invented findings do the opposite, for three reasons.

They are plausible by construction. The model produces what a finding about your domain tends to sound like. Plausibility is the output, not a coincidence.

They usually match what you already suspected. The model is working from your framing, your questions, your language. What comes back is shaped by what you brought. A finding that confirms your hypothesis raises no alarm.

Fabricated quotes are better written than real ones. Real people say "yeah, I mean, it's, I don't really, sorry, what was the question". Invented participants speak in clean, quotable sentences. If a quote in your synthesis would look good on a slide with no editing, be suspicious of it.

The failure mode is not a nonsense output you would laugh at. It is one theme out of six being ungrounded, in a document where the other five are solid. That is the one that reaches the roadmap.

The habit: every claim traces to a line

There is one discipline that makes AI synthesis safe, and it is not clever prompting. It is refusing to accept any finding that does not come with a specific, locatable piece of evidence, and then actually going and looking for it.

That means asking for output in a form you can check, and it means doing the check. A quote with a participant label and a rough position in the transcript is falsifiable in a few seconds: you search the transcript for a distinctive phrase. Either it is there or it is not. A theme described in the abstract is not checkable at all, which is why the abstract version is the one you must refuse.

❌ Weak prompt

Prompt

Here are 12 user interview transcripts. What are the main themes?

Output

Three themes emerge. 1. Users feel overwhelmed by the number of options at setup. 2. Trust in the pricing is a recurring barrier. 3. Participants want more control over notifications. As one participant put it, 'I just want it to stop shouting at me.'

Every one of these could be true. None of them is checkable. There is no participant label, no count, no position in any transcript, and the quote is suspiciously well phrased. You would have to re-read all twelve transcripts to find out which parts are real, which defeats the purpose.

✅ Good prompt

Prompt

Here are 12 transcripts, each headed P1 to P12. Produce themes under these rules. For every theme: list which participant numbers expressed it and how many of the 12 that is; give two verbatim quotes copied exactly, including hesitations and filler words, each with its participant number and the section of the interview it came from; and add a line called COUNTER listing any participant who said something that contradicts the theme. If you cannot find two verbatim quotes for a theme, do not report the theme. Do not paraphrase inside quotation marks.

Output

Theme 1: setup feels like too many decisions at once. Expressed by P2, P4, P9, P11 (4 of 12). Quote, P4, early section: 'there's like, um, six things it wants me to pick and I haven't even, I don't know what any of them do yet'. Quote, P11, mid section: 'I closed it and came back the next day honestly'. COUNTER: P6 said the options were 'fine, I just clicked through'. Theme 2 was dropped: only one verbatim quote available.

Now every claim is attached to something you can search for. The counts stop a single loud participant becoming a finding, the COUNTER line surfaces the disagreement a tidy summary would smooth away, and the instruction to drop unsupported themes means the model discards rather than embellishes.

Two details in that good prompt do most of the work. Asking for hesitations and filler words is a trap for fabrication: invented speech comes out fluent, and a model asked to preserve "um" in a quote it made up tends to produce something that reads oddly. Asking for a COUNTER line matters even more, because the natural pull of summarisation is towards tidiness, and tidiness is where nuance dies.

Prompt you can copy: the traceability check

For the synthesis you just produced, go back through it claim by claim.

For every quotation you used, copy the surrounding three lines from the original transcript I gave you, so I can see it in context.

For every theme, list the participant numbers again and state how you decided that participant expressed it.

If any quotation or theme cannot be traced to specific text I supplied, say so explicitly and remove it. Do not repair it by finding something similar.

Run that as a separate message, after the synthesis. Asking it to be careful in advance is weaker than asking it to audit work it has already committed to.

Checkpoint

A model will invent themes and quotes that read exactly like the real ones, so accept no finding that does not come with a verbatim quote and a participant you can trace back to the transcript.

Other things worth asking for

Once traceability is in place, a few more requests earn their time.

Ask what is absent. "What did I never ask about that these transcripts suggest matters?" This is one of the better uses of the tool, because it is a question about the shape of the material rather than a claim about the world.

Ask for the minority view. "Which observations came from only one participant?" Those are not findings, but they are often the most interesting thing in the pile.

Ask it to argue the opposite. "Here is my conclusion. Using only the transcripts, build the strongest case against it." You will find out quickly whether the evidence is as one-sided as you thought.

Work in one sitting per study. Long conversations drift, and a drifting session starts treating its own earlier summary as the source. If you want the mechanics of that, Context Engineering covers it properly.

Read the transcripts yourself first, even quickly. Synthesis you cannot sanity check is synthesis you cannot defend in a room, and someone will eventually ask you which participant said that.

Checkpoint

Beyond checking quotes, the highest value questions are about absence, disagreement and the minority view, because those resist the pull towards a tidy summary.

📝 Quiz

Question 1 of 4

Why is a beautifully phrased quote in an AI synthesis a warning sign?

Found this useful? Pass it on.