Mira - AI moderated interviews that read how people feel
by•
Unlike AI tools that stop at interview + transcript, Mira is a full AI researcher — plans studies, recruits globally (100M+ panel, 120 countries), runs dynamic interviews with intelligent probing, and uniquely captures what participants say AND feel via real-time facial coding, voice emotion AI, and webcam eye tracking.
Extracts themes, generates insights, and produces research reports automatically. 17 patents. 70+ languages. Trusted by Unilever, Nestlé and 150+ global brands. $25M Series B.


Replies
Per-frame confidence scoring is the right instinct. The bit I'd push on at 120-country scale is cross-cultural validity of the facial and voice layer. Most action-unit and voice-emotion models train on largely Western data, and expression-to-affect doesn't transfer cleanly: gaze aversion, smile intensity, vocal pitch carry different meaning across cultures, so a 'how they feel' score can be confidently miscalibrated for a Jakarta panel while looking fine on a London one. Do you re-validate the emotion mapping per region, or is it one global model?
Mira
@dipankar_sarkar Dipankar, this is one of the most precise critiques anyone has landed today, and it deserves a direct answer rather than a deflection.
The short version: we don't run a single global model applied uniformly. But we also won't claim full per-region revalidation across all 120 countries; that would be overstating where the field currently stands.
Here's what we actually do:
Individual baseline calibration is the first layer. The model doesn't score against a universal affect norm. It measures deviation from each participant's own established neutral. That sidesteps the bulk of cross-cultural expression-to-affect transfer error; a Jakarta participant's gaze aversion is measured against their own baseline, not against a Western norm.
Our training data significantly spans non-Western populations. Nine years of data collection across APAC, MENA, South Asia, and Latin America. The data collection platform uses inter-rater reliability thresholds before any tag enters the training set, and regional annotators are involved in the process.
Multimodal fusion reduces single-signal miscalibration risk. Vocal pitch miscalibrated for a specific region still gets weighted against facial and gaze signals. A confidently wrong facial read is harder to sustain when voice and attention data disagree with it.
Where we're still building: complete per-region model variants at the AU level. We surface confidence scores precisely because of this — a researcher in a less-validated market should weight signals accordingly.
If you'd like to go into architecture at the model level, we would like to speak to you more on this → https://www.entropik.io/book-demo
How does the facial coding and emotion AI handle privacy and consent across different regions, especially with GDPR and other strict regulations?
Mira
@ufukaktugba
Hi Ufuk, great question.
Consent is handled per session: participants see a clear explanation of what data will be captured (facial, voice, eye) before any session begins and must actively opt in.
On GDPR: we are fully GDPR compliant. Data is processed within compliant infrastructure, participants can request deletion, and researchers control retention policies. We also have SOC 2 Type II certification.
For APAC markets we follow regional equivalents (PDPA in Singapore, PIPL in China). If you are evaluating for a specific region, happy to go deeper.
BetterClaw
This is a really interesting take on AI-powered research.👏🏻
Going beyond transcripts to capture emotions and behavioral signals could unlock much richer insights for product teams. Curious, how do you balance those advanced features with participant privacy and consent during interviews?
Mira
@worksforme Hi Laiba — the balance is built into the session flow itself. Every participant gets a clear explanation of what will be captured before they join — webcam access for facial and eye data, microphone for voice emotion. They can decline any of these and still participate, with the system adapting to use only the available signals.
No facial recognition is used anywhere — we only measure expressions, not identity. Data is never sold or used for external training without explicit permission. Researchers control retention and deletion.
Happy to walk through the full consent flow if that would help.
I have to confess I was worried at first because I read "interview" and assumed it was another AI recruitment solution... which opens a huge debate about ethical recruitment and the legislation around automated decision making. However, digging deeper I realise this is actually about user experience feedback... which I guess means there is no reason not to try and capture the "give aways" in terms of facial expressions and pauses etc. There has been some work done around this in the recruitment space (I remember one video platform differentiating between whether a candidate looked to the left or right when answering because it suggested whether they were using the creative or logical part of their brain... as above I think this has now been outlawed in HR processes). But in your space, I think you have the opportunity to get creative and you certainly seem to have done some great work. Best of luck to you.
Mira
@martin_tanner Martin, this is a genuinely important distinction you have drawn, and you are right on both counts.
Facial coding in recruitment is rightly controversial and in many jurisdictions rightly restricted. The core ethical problem is using biometric signals to make high-stakes decisions about people; employment, credit, access, without their meaningful understanding or control. That is a fundamentally different situation from what we do.
In consumer and user research, participants choose to take part, they can stop at any time, and the output of the research does not affect them, it affects product decisions. The emotional signals Mira captures are used to help brands understand how people genuinely respond to products and concepts. No decision is made about them as individuals.
We also do not use facial recognition; we measure expressions, not identity. Participants are anonymous at the analysis level.
You are right that the research space is where this can be done responsibly. That is why we built here. Appreciate you thinking it through properly rather than reacting to the phrase "facial coding."
How does the facial coding and emotion AI actually perform across different cultures since expressions and emotional cues vary so widely globally?
Mira
@zgreg0z Our models are trained across demographic groups rather than on a single population, and we apply normalisation based on cultural benchmarks we have collected to reduce bias in emotional scoring. The goal is to measure genuine reactions, not to apply one cultural standard globally. We also let researchers set study-specific calibration. Emotion expression does vary across cultures and we do not claim perfect universality — it is an ongoing area of work. For global studies we recommend always combining the emotional signal with the transcript and not reading it in isolation.
The facial coding and emotion AI pieces genuinely surprised me, that kind of read on participants usually takes a trained researcher. Watched it pull themes from a test interview in minutes.
Mira
@cal_turkan32795 "That kind of read usually takes a trained researcher" is the most accurate description of what we were trying to automate. The themes in minutes is the time compression — the depth of signal you are working from stays the same.
How does the facial coding and eye tracking actually work in practice — is it through a browser plugin, a mobile app, or do participants need special hardware?
Mira
@nkurumak28108 No plugin, no app, no special hardware. Everything runs in the browser; participants just grant access to the webcam and microphone when they join. There is a quick calibration check before the session starts to make sure lighting and camera positioning are good. Works on any standard laptop or desktop webcam.
How does the facial coding and eye tracking piece actually work in practice, do participants need to opt in and run any special setup on their end?
Mira
@ezelsukut Everything runs in the browser, no downloads or setup. Participants grant camera access, run a short calibration (only for eye tracking), and the session starts on a standard laptop or phone with in-built cam. Fully consent driven.
I have run plenty of user interviews by hand, so the moderation part I get. The reads-how-people-feel part is where I would love more detail: inferring emotion from voice or wording is powerful, but it is also the kind of signal that can mislead a decision (someone nervous is not someone negative). How do you present that layer to the researcher, as a hint to probe further or as scored data? The difference feels important.
Mira
@virko_kask Virko, you've named the exact distinction that separates useful emotion data from noise — and it shaped how we built the output layer.
The short answer: it's a hint to probe further, not a verdict.
We made a deliberate decision not to present emotion as scored data; the researcher is supposed to act on it directly. A nervousness signal doesn't get labeled "negative response." It gets flagged as a moment worth returning to, a timestamp when verbal and non-verbal signals diverged or when an emotion was sustained long enough to be meaningful.
In the report, it looks like a highlighted clip with the emotional trace underneath it. The researcher sees: what was said, what the face and voice were doing at that moment, and how long it lasted. No single-number sentiment score. No resolved verdict. The interpretation is yours.
Where we do add structure, we distinguish between transient signals (a quick flash of surprise, a one-second hesitation) and sustained patterns (3–5 seconds of consistent emotional signal across modalities). Transient signals are surfaced as context. Sustained patterns are what get elevated as findings. That threshold matters for exactly the reason you named, nervousness is not negativity, but sustained disengagement during a key product moment is worth a conversation.
The emotional layer is evidence, not a conclusion. Would love to show you what this looks like on a real study, especially with your interview background. I think you'd spot things most researchers miss. Happy to set up a call → https://www.entropik.io/book-demo
@mridhu_varshini_ Thank you for the detailed answer, this is genuinely thoughtful design. The transient vs sustained threshold is the right cut, and surfacing the divergence as a clip with the trace under it, instead of a number, is exactly the difference between evidence and verdict. I am heads-down on my own launch for the next two weeks, but I would honestly enjoy seeing this on a real study after that. Good luck with the launch, Mira deserves the attention it is getting.
Mira
@virko_kask Thank you, Virko, this genuinely made my day :) The way you framed it, "evidence not verdict," is exactly the distinction we spent a long time trying to get right in the design. It's validating to hear it land that way from someone who's run as many interviews as you have.
And yes!!!! Mira deserves this attention, and honestly, so do you for asking the questions that pushed us to articulate it properly. This launch has been our killer ship moment, years of building, finally in front of the right people.
Two weeks go fast. Ping me when you're out the other side, and we'll get something on the calendar. I genuinely think you'll have interesting observations after seeing it in a real study.
The mid-interview probe on the say/feel mismatch is the part that interests me. I run the cheap cousin of this for my own app, a panel of simulated user personas that scores LLM output before a change ships, and the one thing simulation can't give me is exactly that hesitation signal you're reading off real faces. When facial coding and the transcript disagree, which one do your reports trust? I'd want the raw disagreement surfaced, not resolved for me.
Mira
@narek_keshishyan Narek, the fact that you've already built a simulated persona scoring layer tells me you'll get more out of this than most.
To your specific question: the disagreement is surfaced, not resolved. That's intentional and non-negotiable for us.
When facial coding and transcript diverge, someone says "yes, that makes sense" while their face reads confusion and their voice drops engagement — the report shows both signals side by side with timestamps. The researcher sees the verbal response and the emotional trace underneath it. We don't collapse them into a single confidence score or pick a winner. That would defeat the entire point.
What the report flags is the gap itself, the moment where the two signals split. You get the clip, the transcript line, the emotion curve, and a marker that says "these disagreed here." What it means is yours to interpret.
Where we do take a position: we surface which signal is sustained longer. A one-second facial flicker during a considered verbal answer is weighted differently from a 6-second emotional hold that contradicts a polished verbal close. But that weighting metadata is visible, not hidden.
Your simulated personas are great for pre-ship scoring, but the hesitation signal on a real face during a real task is a different class of data. I'd genuinely love to walk you through how the disagreement layer looks in a real study. Can we set up a call? → https://www.entropik.io/book-demo
@mridhu_varshini_ surfaced not resolved is the answer I was hoping for. weighting by how long a signal is sustained makes sense too, a six-second hold is a different animal from a flicker. no call promises mid-launch-week, but I'll poke at the demo materials. appreciate you writing this out properly.
Mira
@narek_keshishyan Glad it landed that way. "surfaced not resolved" is genuinely the design philosophy, not just a talking point.
No pressure on the call. When you do poke at the demo materials, if anything raises a question worth going deeper on, I'm here. Would be curious to hear what you think after you've had a look.