The say/feel gap in research; how big is the problem actually, and what do you do about it?
There's a well-documented phenomenon in consumer and user research that most practitioners know intuitively but rarely name directly: people don't report their experience accurately.
Not because they're dishonest. Because self-reporting is hard. In the moment of an interview, participants are performing a version of themselves. They round off hesitation. They describe their behaviour more charitably than it actually was. They say "yes, I'd probably use this" when what they felt was closer to "maybe, under the right circumstances, if the price were different." The social pressure of being in a conversation, even with an AI, shapes what gets said.
This is what we call the Say-Do Gap. The distance between what someone tells you and what they actually feel or do.
It shows up everywhere. In concept tests where participants say they love a product they'd never buy. In usability sessions where someone says "this makes sense" while visibly struggling. In brand perception studies where stated attitudes don't match purchasing behaviour.
The traditional workarounds are imperfect. Projective techniques add noise. Implicit association tests are hard to run at scale. Follow-up probing helps but depends entirely on the moderator catching the right moments. And even the best human moderator misses things, they can't track vocal confidence, facial micro-expressions, and conversation content simultaneously across 40 interviews.
At Decode, the reason we built an emotion layer into Mira wasn't to replace researcher judgement. It was specifically to give researchers a way to see the moments where what someone said and what they felt came apart, and to surface those moments as evidence, not as a resolved verdict.
I'm curious how others in this room handle it.
Do you design studies specifically to work around the say/do gap?
Do you treat self-reported data with a standard discount?
Or have you found methods that get closer to what people actually feel?


Replies
@productrambler Does it ever say i don't know instead of guessing?
Mira
@lee_jay1 It should, and in practice, it does through confidence scoring and frame-level flagging. Low-quality frames (poor lighting, occlusion, atypical expression) are flagged as low-confidence rather than interpolated. The system won't produce a clean emotional read on a frame it can't read cleanly. That matters because a confident wrong answer is worse than a flagged uncertainty in a research context.
@productrambler This is a very real problem we see in moderated tests. From being a target of attention or feeling that users are themselves on trial, it's difficult to gauge the response.
Some of the methods we tried over the years.
- Look at nonverbal cues: how do they react, what do they do with their hands, etc.?
- On purpose we stage some easy tasks at the start of the session, to build trust and confidence
- We mix the questions, change order, run different sets etc.
Traditional user tests work well, but I think they don't reflect real users. A/B tests on a fraction of the real audience for specific journeys or cohorts lead to much better results.
Mira
@roopesh_donde These are exactly the right instincts, and it's worth noting that each of them is a workaround for the same underlying constraint: a human moderator can only track so many signals simultaneously across so many sessions.
The staged easy tasks point is particularly sharp. Building trust early changes participant behaviour for the rest of the session. That's a design principle we've tried to replicate in how Mira structures the opening of an interview, lower stakes questions first, before anything emotionally loaded.
The A/B test point is honest and important. For decision-level validation, behavioural data on real users beats stated preference every time. Qual research is best used to understand why the A/B result happened, not to predict it.
I live one domain over (hiring screens rather than research interviews) and the say/do gap is the whole story there too: a resume is a performed self, and AI writing tools just made the performance flawless. The working principle that survived for me sounds exactly like your evidence-not-verdict framing: an inferred signal (emotion, confidence, AI-likeness) is only allowed to point a human at a moment worth probing, never to carry a decision by itself. Two reasons. Inferred signals misfire on the wrong people: a nervous voice is not a negative answer, a non-native writing style is not a template. And decisions made on unverifiable signals do not survive the first challenge from a stakeholder. Evidence the human can check, judgment the human keeps. The gap does not close, but it stops being invisible, and that is most of the value. One question back: across your 40-interview example, how do you handle the temptation to just sort by the emotion layer and only watch the flagged clips? That feels like the quiet path back to verdict-by-score.
Mira
@virko_kask This is the sharpest question in the thread and I want to answer it honestly.
The temptation is real. Flagged clips are faster. Sorting by emotion signal is convenient. And if you're under time pressure, it's easy to treat the flagged moments as the findings rather than as the starting points.
We've tried to design against it structurally. The default report view shows the full interview timeline first, with emotional markers overlaid, not a sorted list of high-signal clips. You have to actively choose to filter by emotion signal, and when you do, the surrounding context (what came before and after) stays visible. We also surface low-emotion moments alongside high-emotion ones, because a sustained neutral during a moment you expected to generate a reaction is itself a signal worth investigating.
But design can only do so much. The real answer is that the risk you're naming, verdict-by-score through the back door, is a research literacy problem as much as a tool design problem. Researchers need to be trained to treat flagged clips as hypotheses to interrogate, not findings to report. That's a harder problem than any UX fix.
Your framing, "inferred signals can only point a human at a moment worth probing, never carry a decision by itself" — is exactly the principle. The challenge is making sure that principle survives contact with a deadline and a 40-interview dataset.
@productrambler Design against the default, then train for the rest, that is a fair split of the burden. The full-timeline-first default and surfacing expected-but-absent reactions are both genuinely good calls, the second one especially: silence where you expected a spike is the finding nobody goes looking for. And agreed, the last mile is literacy under deadline pressure, no UX can carry that alone. Thanks for engaging with the question this seriously, it shows in how the product is put together. Good luck with the launch week.
Love seeing more attention on the difference between stated preferences and real behavior. That's where many product decisions succeed or fail.