The Interview Survives: Brendan Jarvis on What AI Still Cannot See

The Interview Survives: Brendan Jarvis on What AI Still Cannot See

Brendan Jarvis discussing AI, synthetic users and the future of user interviews for DesignWhine Issue #20
Brendan Jarvis on AI user research, synthetic users and why real user interviews still matter.

There is a particular kind of confidence that AI is very good at producing. A clean transcript. A tidy cluster of themes. A polished findings deck. A plausible answer from a synthetic participant that sounds exactly like something a real customer might say.

For Brendan Jarvis, that fluency is precisely where the danger begins.

Jarvis is an independent UX researcher and the Managing Founder of The Space InBetween, a specialist UX research practice in New Zealand. He also hosted the Brave UX podcast, where he spent years speaking with researchers, designers and product leaders about the judgment behind good practice, not just the methods around it.

This is not DesignWhine’s first conversation with Jarvis. We interviewed Brendan earlier about the Brave UX podcast, speaking with him about bravery in product work, assumptions about usability and value, and the thinking behind better UX practice. This new conversation continues that relationship in a very different moment, as AI begins to reshape the research process itself.

That distinction matters in the current AI moment. Across user research, tools are moving into planning, recruitment, moderation, transcription, synthesis and reporting. Some of this is already genuinely useful. Some of it looks more capable than it is. And some of it may be quietly changing what organisations are willing to count as research.

For DesignWhine’s Issue #20, The Last User Interview, we asked Jarvis where AI is making research better, where it is merely making research look better, and what happens when plausibility starts to stand in for lived experience.

The Easy Wins, and the Harder Problem

Jarvis does not reject AI in research. In fact, he starts with the areas where the gains are already obvious.

“Transcription is the most obvious improvement. It is fast and accurate enough. Automated note-taking and initial tagging are useful too, as long as someone still takes the time to read the raw material, and it can be tempting not to.

Synthesis is where I would be most careful. The tools I’ve seen are very good at producing something that looks like a set of themes. Plausible themes and accurate themes are not the same thing though, and the tools can miss outliers. In my experience those outliers are often the insights.

An area few people talk about is recruitment, where AI may be making it harder. Screeners are actively being gamed by people using AI. If you are paying participants, that is a problem.”

“Plausible themes and accurate themes are not the same thing though, and the tools can miss outliers. In my experience those outliers are often the insights.”

The concern echoes a broader tension we found while examining synthetic users versus real users: AI is often strongest where the task is structured, repeatable or linguistic, and weakest where the research depends on context, surprise and interpretation. Nielsen Norman Group has similarly argued that AI can accelerate parts of planning and analysis while remaining a poor fit for messy, semi-structured research where the moderator must change direction in response to what happens in the room.

What Should Never Become a Shortcut

The temptation, of course, is not merely to use AI for support work. It is to let the support system inherit the judgment.

Jarvis draws the boundary around three things.

“Three things come to mind. Deciding what the research is for. Reading the output critically instead of accepting it. And moderation.

On moderation, the case for handing it over is that AI can run a hundred sessions while you sleep. The case against is that moderation is where you decide to abandon the plan. An AI moderator will work through the guide. A researcher notices the pause, hears the contradiction between what someone just said and what they did two minutes earlier, and improvises the next question. That judgment is the value.”

That is also where the efficiency story becomes complicated. AI moderation can clearly expand scale. Recent Nielsen Norman Group testing of AI-moderated interviews found useful cases for structured feedback, screening and multilingual research, while still concluding that they do not replace the depth of expert-led semi-structured interviews. Scale and depth are not interchangeable outcomes.

Plausibility Is Not Experience

If there is one word that keeps resurfacing in Jarvis’s answers, it is plausible. Synthetic users are compelling because they produce answers that sound right. But sounding right and being evidence are two very different things.

“An AI moderator will work through the guide. A researcher notices the pause, hears the contradiction between what someone just said and what they did two minutes earlier, and improvises the next question. That judgment is the value.”

“We have learned what a plausible answer looks like, which can be useful for pressure-testing a discussion guide or anticipating objections before you present findings. But it tells you about the distribution of language on a topic, not about any person.

This plausibility can be a problem. These models are built to produce the expected answer. As researchers our value is tied to our ability to find the unexpected one, the thing nobody predicted. Are we looking at an insight, or confirmation of our own bias?”

That question sits at the centre of the current synthetic-research boom. In our review of Articos, we found much the same tension: AI-generated research can be remarkably fast and coherent, but coherence itself is not validation. The more polished the output becomes, the easier it may be to forget how uncertain the evidence underneath it still is.

When Weak Research Starts Looking Strong

The organisational pressure behind all of this is obvious. Research is expensive. Recruiting takes time. Moderation takes time. Analysis takes time. AI offers a version of the process that is faster, cheaper and easier to scale.

Jarvis worries less about organisations explicitly deciding that synthetic users are superior to real users than about something more subtle: losing the ability to see when the substitution has failed.

“Yes, and the risk is not that organisations choose synthetic over real. It is that they cannot tell when the choice has cost them anything.

A weak study with real users looks weak. Mismatched recruits, thin evidence, poorly framed insights. A weak synthetic study reads beautifully. There is no error signal, so the budget for real contact gets cut and the cost does not show up for several quarters.

On democratisation, what we’ve actually democratised is the production of research artifacts, not research judgment. Yes, more people can make a great-looking findings deck, but how many of them are changing what gets built for the better?”

The Interview Does Not Disappear

Our issue is called The Last User Interview, deliberately provocative language for a field currently being told that simulated participants, autonomous agents and AI moderators can absorb more of the work that once required direct contact with people.

Jarvis does not think the interview is going away. He thinks its place in the research stack may change.

“The good news is the interview survives, and so does watching someone use the thing, because of what is not said.

In a synthetic user session, there is no user in the room. There is no pause before an answer, no scrolling back to reread, no hand that goes to the button and comes away again, no “yeah, it’s fine” when the opposite is meant. None of that is being hidden from the model. It does not exist. Those are the moments that make us ask the kind of question that opens the whole session up.

“A weak synthetic study reads beautifully. There is no error signal, so the budget for real contact gets cut and the cost does not show up for several quarters.”

Where I see these tools becoming powerful is in heuristic work. Point an agent at a build with a defined set of heuristics or accessibility criteria and it will take the heavy lifting out of the early work. It does not get bored on item two hundred. The rules are explicit and the results are checkable. But that is expert review, not user research.

An agent does not feel anything about the outcome, and more to the point it has nothing riding on it. It is not worried the button will charge its card, it is not doing this on a phone in a carpark with one bar, and it has not already tried twice this week. You could give an agent a bad connection and a history of failed attempts. You cannot give it a stake in what happens next. Emotion and context are not features that get added in a later release. They come from being a person in a situation.

So no, I do not think we run out of interviews. I think they stop being the default and become the thing we reach for when the cheaper AI-led methods cannot answer the question. That is likely healthier than what we do now, as long as we check the cheaper methods.”

That may be the more interesting future than either side of the current debate tends to offer. AI does not have to replace user research to change it. It may instead force teams to become more deliberate about why they are speaking to users in the first place, which questions deserve direct human contact, and which parts of the process can safely be delegated.

If Jarvis is right, the interview survives not because the industry refuses automation, but because there are still moments in research where the signal is inseparable from the person producing it. A pause. A contradiction. A hand hovering over a button. A stake in what happens next.

Share this in your network
retro
Written by
DesignWhine Editorial Team
Leave a comment

1 Comment
  • Where do you draw the line between AI-assisted research and the point where human judgment becomes essential? I’m especially curious whether you’d trust AI more for synthesis, moderation, recruitment, or heuristic evaluation.