How I synthesize user interviews with Claude — without letting it invent the insights
There's a lazy version of this that a lot of designers are already doing: paste the transcripts, ask for “the key themes,” copy the bullet points into the readout. It produces something that looks like research and isn't. The model gives you a tidy summary of what people talked about, everyone nods, and no design decision actually changes.
I use Claude on interview data constantly, but not like that. The difference between a summary and an insight is the whole job. Claude will happily blur that difference unless the process protects it.
Quotes, themes, insights — three different things
Getting these straight is most of the battle.
A theme is a label for what people talked about: “users mentioned onboarding confusion.” A quote is the evidence: “I didn't know where to start, there were too many options.” An insight is a non-obvious conclusion that explains behavior and implies a decision: “first-time users read feature richness as complexity — the product's biggest selling point becomes its main abandonment trigger.”
Most AI-assisted synthesis stops at themes and calls it a day. Themes are the cheapest output and the least useful. A real insight has to name a mechanism, survive the “so what” test, and change how you'd frame the design problem. If it doesn't do all three, it's still a summary.
The four stages
I run interview data through four passes, and the order matters. Each one has a checkpoint where my judgment shapes the output before it hardens into something I'd have to defend later.
Stage 1 — Preprocessing. Structure the raw data before any analysis. This is the step everyone skips and it's the one that protects you. Uneven inputs — one detailed interview, three thin ones — make the model over-weight the rich transcript and hallucinate patterns from the sparse ones. So I have it extract, per interview, who the person is, the specific moment adoption broke down, any emotion or social dynamic mentioned, and the direct quotes worth keeping. Verbatim quotes stay verbatim when someone's exact words reveal why they decided something. Everything neutral gets paraphrased. The instruction I actually use: “Don't analyze — just structure what's in the data. If an interview has limited detail, flag it rather than filling the gaps.”
Stage 2 — Pattern detection. Now find where different people describe the same underlying experience in different words. The key word is underlying. I tell it to group by shared experience, not shared topic — and to protect outliers. An experience that shows up once but reveals a mechanism is worth more than a theme that appears three times. Contradictions get flagged, not smoothed over. Repetition is the obvious signal, but it isn't the only one.
Stage 3 — Insight sharpening. Convert patterns into insights by naming the mechanism — why the behavior was the rational thing to do from the user's point of view. Each insight has to carry three things: the mechanism (why this made sense given their situation), the “so what” (what it means for design direction, not a solution yet), and the riskiest assumption (what would have to be true for this to hold that the data doesn't fully prove). That last one doubles as your next research question, already written.
The trap in this stage is asking “how do we fix this” too early. Prompt for solutions before the mechanism is clear and you get generic UX recommendations that would fit any product on earth.
Stage 4 — Implication mapping. Only now do I connect insights to decisions — and I ask for a decision space, not a single answer. For each insight: two to four design moves that address the mechanism rather than the surface symptom, including what to stop doing, ranked by how directly they hit the mechanism, each tagged with the assumption it depends on. That gives me something to test, not something to blindly build.
The move that separates senior from junior
Two prompting habits do most of the work.
The first is asking what was not said. “No participant mentioned price, despite this being a consumer product” is a finding, not a gap. The model is surprisingly good at flagging conspicuous absence, and absence is exactly what a tired researcher stops noticing around interview four.
The second is the mechanism question: when behavior looks irrational on the surface, ask what would make it the most logical thing that person could do given their situation. That single reframe is where the real insights live. Champions going silent on a B2B rollout turned out to be rational reputation management — they protect themselves politically when adoption stalls. Users who don't come back after a disruption aren't churning; nothing re-invited them. The behavior only makes sense once you find the mechanism, and the mechanism is what reframes the design problem.
That last example taught me the reframe I keep coming back to: B2B adoption is a social and organizational event, not a product-experience event. Stop at “onboarding is confusing” and you'll ship something competent that solves the wrong problem.
What the process actually buys me
Not speed, mostly. What it buys is a visible reasoning chain I can defend to stakeholders, checkpoints where I catch a bad insight before it reaches the readout, design briefs specific to what the data actually said, and a built-in next research question for every conclusion.
The role instruction matters more than people think, too. “UX researcher” gives you generic output. “Behavioral researcher analyzing why a rational person would make this choice” activates a genuinely more useful frame. And I cap the output — “as many insights as you find” produces padding; a hard maximum of three forces prioritization and makes me choose.
The honest summary: Claude is a very good thinking partner and a very bad researcher. It will find patterns all day. Whether those patterns are insights or just tidy summaries is still my call — and the four stages exist so I never hand that call to the model by accident.