Mira - AI moderated interviews that read how people feel

by
Unlike AI tools that stop at interview + transcript, Mira is a full AI researcher — plans studies, recruits globally (100M+ panel, 120 countries), runs dynamic interviews with intelligent probing, and uniquely captures what participants say AND feel via real-time facial coding, voice emotion AI, and webcam eye tracking. Extracts themes, generates insights, and produces research reports automatically. 17 patents. 70+ languages. Trusted by Unilever, Nestlé and 150+ global brands. $25M Series B.

Add a comment

Replies

Best

Congratulations on launch! A lot of the questions are about accuracy and privacy, so I'll ask a different one: how does the emotion reading hold up across cultures and languages? People show feelings differently depending on background, and my audience is fairly reserved by nature. Does the model account for that, or is it mostly calibrated to more expressive participants?

 Anastasiia, this is one of the most important questions anyone can ask about emotion AI — and one we've had to answer in practice across 150+ brands in markets where emotional expressiveness varies significantly.

A few things we do specifically for this:

Individual baseline calibration is not a universal norm. At the start of every session, Mira establishes a neutral baseline for that participant. What matters is deviation from their baseline, not deviation from an average. A reserved participant who shifts even slightly is flagged because for them, that shift is meaningful. You're not comparing them to a participant in a different country who naturally gestures more.

Multimodal cross-referencing helps here significantly. Voice tone and pacing convey emotional signals across cultures in ways that facial expressions alone don't. Someone from a more reserved cultural background who shows minimal facial movement will often convey more signals in their voice, hesitation, changes in pacing, or a drop in confidence. Mira reads both simultaneously.

We've trained specifically on diverse participant pools. Our models have been built on research across APAC, MENA, Europe, and the Americas, not just Western expressive populations. That said, we're transparent: some cultural nuances remain an active area for improvement, and we'd rather flag lower-confidence signals than assert false precision.

The output accounts for this contextually. Signals are always shown with confidence scores and duration markers. A researcher working with a reserved audience can set their own threshold for what's worth acting on.

If your audience is particularly reserved by nature, I'd love to show you a session with that kind of participant profile, specifically, it's actually where the multimodal approach shows its biggest advantage over facial-only tools.

Worth a call?

 thank you for taking the time to explain this so fully. The baseline calibration and the multimodal reading for reserved participants answer my concern well. I'll keep Mira in mind for when I'm at that research stage. 

 , sure. We can reconnect anytime when the research stage comes in. Happy to be connected :)

the facial coding during interviews genuinely caught me off guard, you can actually see the moment a respondent's expression shifts on a tricky question and it ties back to the insight automatically.

 The "ties back to the insight automatically" is the AI Copilot connecting the emotional moment to the theme it contributed to. You are not just seeing a flag; you are seeing why it mattered in the context of the full study.

How does the facial coding and eye tracking actually work in practice with participants who are on their phones or in different lighting conditions? Curious how reliable that data really is across such a massive global panel.

 Great question Miraç.

A few specifics on how we handle this:

Phones: Mira currently runs in-browser on desktop, participants join via a standard webcam, not a mobile camera. We found controlled desktop sessions produce more reliable facial signal than mobile, where camera angle and movement vary too much. For global panel recruitment, participants are briefed on the setup requirement before joining.

Lighting: Before every session, participants go through a quick calibration check, the system guides them to adjust their face position and lighting until confidence scores are acceptable. Frames that fall below our threshold are automatically excluded from the emotional signal rather than being interpolated. We never guess on weak data.

Global reliability: Our models are trained across demographic groups and we apply cultural normalisation benchmarks to reduce bias in emotional scoring. The signal is most reliable when used comparatively, patterns across a study group, rather than reading single participant moments in isolation.

Honest caveat: webcam-based tracking is less precise than lab infrared. For questions like where attention went first or whether someone showed hesitation at a specific moment, it works well. For pixel-level gaze accuracy, a lab setup still wins.

Happy to go deeper on any of this.

When Mira spots a say/feel mismatch and digs deeper on its own, how do you keep that follow-up from leading the participant? A moderator reacting to visible confusion can easily plant the doubt rather than uncover it. Is the probe neutral by design, or tuned per study?

 David, this is the question we lost sleep over. A leading probe is almost worse than no probe — it plants the narrative.

A few design decisions we made:

Probes are behaviorally triggered but linguistically neutral. When Mira detects a signal — a hesitation, an emotional shift, a response that doesn't match the facial/voice pattern — it doesn't say "you seemed confused." It says something like "You paused there — tell me more about what was going through your mind." The trigger is emotional, the language is open.

No interpretive language in the probe. Mira never names the emotion it detected back to the participant. It asks outward, not inward. This is a deliberate constraint built into the probing engine.

Tunable per study type, not per participant. Concept testing probes are calibrated differently from usability or brand research — because the type of signal differs. But within a study, every participant gets the same probe structure. That keeps cross-participant comparisons valid.

Probe depth is capped. Max 2 levels of follow-up on any one signal — so the interview doesn't become an interrogation.

The honest caveat: no probe is perfectly neutral. But compared to human moderators who vary tone and word choice across 40 interviews — Mira is at least consistently neutral. Happy to show you a live session →

What stands out is the integration of facial coding and voice emotion AI directly into the interview flow rather than bolting them on as an afterthought. That feels like a real research instrument, not just a chat wrapper dressed up with sentiment scores.

 Ayşe, thank you — "research instrument, not a chat wrapper" is exactly the bar we held ourselves to.

The decision to build the emotion layer into the interview flow rather than running it as a post-processing layer was deliberate. When emotion signals are captured in real time during the conversation, Mira can actually respond to them — adjusting its probing when it detects hesitation or conflict between what someone says and how they react. That closed loop is what makes it feel like a moderator, not a recorder.

The underlying models (facial coding, voice tone, eye tracking) are the same tech we've been building for 9 years across 150+ brand research programs — we just finally wired them into the interview itself.

Would love to show you what a full session looks like end to end →

"Reads how people feel" is the interesting (and risky) part. When it detects hesitation mid-interview, does it adapt its questioning in the moment or just annotate for the researcher afterward? I'm running beta-user interviews right now and what I always miss is what people didn't say, I would love to know if you surface that.

 Anthony, you've named the exact thing most AI interview tools get wrong: they annotate after, when the moment is already gone.

Mira adapts in the moment. When it detects hesitation, an emotional shift, or a mismatch between what's said and how it's said, it probes during the conversation rather than as a post-hoc tag. "You paused there. Tell me more about what was going through your mind." The follow-up happens while the participant is still in the thought, not after they've moved on.

On what people didn't say: this is where the multimodal layer earns its place. Mira flags moments when vocal confidence dropped without a corresponding verbal signal, when attention shifted during a key question, and when facial affect changed but the spoken answer was flat. Those are the gaps, the things a transcript alone will never show you. They surface in the report as flagged moments with timestamps, so you can go directly to "here's where something was happening that wasn't being said."

For beta-user interviews specifically, that layer tends to be most valuable on product moments — the exact second someone's expression changes when they hit a friction point, even while they're still saying "yeah, this makes sense."

Would love to show you what this looks like on a real session, especially with beta interviews, the signal density is usually high →

How does the facial coding and eye tracking actually work on participants who don't have webcams or who join from mobile devices?

 Right now, you would need camera access to have face emotion measurement or eye gaze tracking done. However, Mira also has the ability to measure emotions purely based on tone of the voice and also based on the words that are being used.

How do you prevent Emotion AI from overinterpreting facial expressions or cultural differences in user reactions?

 We avoid cultural overfitting by rejecting static global priors and instead computing affect as temporal deltas against a per user baseline. Our attention layer gates visual action units against orthogonal acoustic and semantic vectors down-weighting isolated facial anomalies as noise unless validated across channels or via active verbal probing.

The Say-Do gap framing is sharp, and running recruit → moderate → analyze → report as one agent is the ambitious part. For a team that already has research infra, two setup questions: can I bring my own recruited participants into a Mira study, or is moderation locked to your 100M panel? And can I export the raw session data — transcripts plus the behavioral signals you capture — into our own repo, or does the analysis stay inside Mira's dashboards?

 Good questions, both get asked by every team with existing infra, so let me be precise rather than salesy.

Bring your own participants: yes. Moderation isn't locked to our panel — you can pipe in your own recruited participants and Mira runs the same moderate → analyze pipeline on them. Two ways teams do this depending on setup:

  1. Link-based: we generate a session link per participant, you distribute through your own recruiting/incentive flow, they land in the same moderated session experience.

  2. API-based: if you've got a recruiting/panel system already, you can push participants in via API and we handle scheduling + session invites from there.

On export: yes, and I'll be specific about what "raw" means. CSV export gives you:

  • Full transcripts (turn-by-turn, timestamped)

  • The say-do gap annotations themselves (where the model flagged claimed vs. observed behavior divergence, with the reasoning trace)

Nothing stays trapped in the dashboard-only view. The dashboard is a convenience layer, not the source of truth — the source of truth is the same data you'd get in the export.

Respect for not hand-waving that, most launch threads would have. One thing I'd add: per-frame confidence is the model scoring its own certainty, so it won't catch systematic bias. A model can be high-confidence and wrong the same way across a whole population and never flag it. The only check I trust is human-coded ground truth sampled per region, which is painful to collect. Which regions have you actually validated against local human coders versus carried over from the base model?