Decode by Entropik - Mira, AI moderated interviews that read how people feel
Unlike AI tools that stop at interview + transcript, Mira is a full AI researcher — plans studies, recruits globally (100M+ panel, 120 countries), runs dynamic interviews with intelligent probing, and uniquely captures what participants say AND feel via real-time facial coding, voice emotion AI, and webcam eye tracking.
Extracts themes, generates insights, and produces research reports automatically. 17 patents. 70+ languages. Trusted by Unilever, Nestlé and 150+ global brands. $25M Series B.


Replies
Per-frame confidence scoring is the right instinct. The bit I'd push on at 120-country scale is cross-cultural validity of the facial and voice layer. Most action-unit and voice-emotion models train on largely Western data, and expression-to-affect doesn't transfer cleanly: gaze aversion, smile intensity, vocal pitch carry different meaning across cultures, so a 'how they feel' score can be confidently miscalibrated for a Jakarta panel while looking fine on a London one. Do you re-validate the emotion mapping per region, or is it one global model?
Mira by Decode
@dipankar_sarkar Dipankar, this is one of the most precise critiques anyone has landed today, and it deserves a direct answer rather than a deflection.
The short version: we don't run a single global model applied uniformly. But we also won't claim full per-region revalidation across all 120 countries; that would be overstating where the field currently stands.
Here's what we actually do:
Individual baseline calibration is the first layer. The model doesn't score against a universal affect norm. It measures deviation from each participant's own established neutral. That sidesteps the bulk of cross-cultural expression-to-affect transfer error; a Jakarta participant's gaze aversion is measured against their own baseline, not against a Western norm.
Our training data significantly spans non-Western populations. Nine years of data collection across APAC, MENA, South Asia, and Latin America. The data collection platform uses inter-rater reliability thresholds before any tag enters the training set, and regional annotators are involved in the process.
Multimodal fusion reduces single-signal miscalibration risk. Vocal pitch miscalibrated for a specific region still gets weighted against facial and gaze signals. A confidently wrong facial read is harder to sustain when voice and attention data disagree with it.
Where we're still building: complete per-region model variants at the AU level. We surface confidence scores precisely because of this — a researcher in a less-validated market should weight signals accordingly.
If you'd like to go into architecture at the model level, we would like to speak to you more on this → https://www.entropik.io/book-demo
BetterClaw
This is a really interesting take on AI-powered research.👏🏻
Going beyond transcripts to capture emotions and behavioral signals could unlock much richer insights for product teams. Curious, how do you balance those advanced features with participant privacy and consent during interviews?
Decode by Entropik
@worksforme Hi Laiba — the balance is built into the session flow itself. Every participant gets a clear explanation of what will be captured before they join — webcam access for facial and eye data, microphone for voice emotion. They can decline any of these and still participate, with the system adapting to use only the available signals.
No facial recognition is used anywhere — we only measure expressions, not identity. Data is never sold or used for external training without explicit permission. Researchers control retention and deletion.
Happy to walk through the full consent flow if that would help.
I have to confess I was worried at first because I read "interview" and assumed it was another AI recruitment solution... which opens a huge debate about ethical recruitment and the legislation around automated decision making. However, digging deeper I realise this is actually about user experience feedback... which I guess means there is no reason not to try and capture the "give aways" in terms of facial expressions and pauses etc. There has been some work done around this in the recruitment space (I remember one video platform differentiating between whether a candidate looked to the left or right when answering because it suggested whether they were using the creative or logical part of their brain... as above I think this has now been outlawed in HR processes). But in your space, I think you have the opportunity to get creative and you certainly seem to have done some great work. Best of luck to you.
Decode by Entropik
@martin_tanner Martin, this is a genuinely important distinction you have drawn, and you are right on both counts.
Facial coding in recruitment is rightly controversial and in many jurisdictions rightly restricted. The core ethical problem is using biometric signals to make high-stakes decisions about people; employment, credit, access, without their meaningful understanding or control. That is a fundamentally different situation from what we do.
In consumer and user research, participants choose to take part, they can stop at any time, and the output of the research does not affect them, it affects product decisions. The emotional signals Mira captures are used to help brands understand how people genuinely respond to products and concepts. No decision is made about them as individuals.
We also do not use facial recognition; we measure expressions, not identity. Participants are anonymous at the analysis level.
You are right that the research space is where this can be done responsibly. That is why we built here. Appreciate you thinking it through properly rather than reacting to the phrase "facial coding."
I have run plenty of user interviews by hand, so the moderation part I get. The reads-how-people-feel part is where I would love more detail: inferring emotion from voice or wording is powerful, but it is also the kind of signal that can mislead a decision (someone nervous is not someone negative). How do you present that layer to the researcher, as a hint to probe further or as scored data? The difference feels important.
Mira by Decode
@virko_kask Virko, you've named the exact distinction that separates useful emotion data from noise — and it shaped how we built the output layer.
The short answer: it's a hint to probe further, not a verdict.
We made a deliberate decision not to present emotion as scored data; the researcher is supposed to act on it directly. A nervousness signal doesn't get labeled "negative response." It gets flagged as a moment worth returning to, a timestamp when verbal and non-verbal signals diverged or when an emotion was sustained long enough to be meaningful.
In the report, it looks like a highlighted clip with the emotional trace underneath it. The researcher sees: what was said, what the face and voice were doing at that moment, and how long it lasted. No single-number sentiment score. No resolved verdict. The interpretation is yours.
Where we do add structure, we distinguish between transient signals (a quick flash of surprise, a one-second hesitation) and sustained patterns (3–5 seconds of consistent emotional signal across modalities). Transient signals are surfaced as context. Sustained patterns are what get elevated as findings. That threshold matters for exactly the reason you named, nervousness is not negativity, but sustained disengagement during a key product moment is worth a conversation.
The emotional layer is evidence, not a conclusion. Would love to show you what this looks like on a real study, especially with your interview background. I think you'd spot things most researchers miss. Happy to set up a call → https://www.entropik.io/book-demo
@mridhu_varshini_ Thank you for the detailed answer, this is genuinely thoughtful design. The transient vs sustained threshold is the right cut, and surfacing the divergence as a clip with the trace under it, instead of a number, is exactly the difference between evidence and verdict. I am heads-down on my own launch for the next two weeks, but I would honestly enjoy seeing this on a real study after that. Good luck with the launch, Mira deserves the attention it is getting.
Mira by Decode
@virko_kask Thank you, Virko, this genuinely made my day :) The way you framed it, "evidence not verdict," is exactly the distinction we spent a long time trying to get right in the design. It's validating to hear it land that way from someone who's run as many interviews as you have.
And yes!!!! Mira deserves this attention, and honestly, so do you for asking the questions that pushed us to articulate it properly. This launch has been our killer ship moment, years of building, finally in front of the right people.
Two weeks go fast. Ping me when you're out the other side, and we'll get something on the calendar. I genuinely think you'll have interesting observations after seeing it in a real study.
The mid-interview probe on the say/feel mismatch is the part that interests me. I run the cheap cousin of this for my own app, a panel of simulated user personas that scores LLM output before a change ships, and the one thing simulation can't give me is exactly that hesitation signal you're reading off real faces. When facial coding and the transcript disagree, which one do your reports trust? I'd want the raw disagreement surfaced, not resolved for me.
Mira by Decode
@narek_keshishyan Narek, the fact that you've already built a simulated persona scoring layer tells me you'll get more out of this than most.
To your specific question: the disagreement is surfaced, not resolved. That's intentional and non-negotiable for us.
When facial coding and transcript diverge, someone says "yes, that makes sense" while their face reads confusion and their voice drops engagement — the report shows both signals side by side with timestamps. The researcher sees the verbal response and the emotional trace underneath it. We don't collapse them into a single confidence score or pick a winner. That would defeat the entire point.
What the report flags is the gap itself, the moment where the two signals split. You get the clip, the transcript line, the emotion curve, and a marker that says "these disagreed here." What it means is yours to interpret.
Where we do take a position: we surface which signal is sustained longer. A one-second facial flicker during a considered verbal answer is weighted differently from a 6-second emotional hold that contradicts a polished verbal close. But that weighting metadata is visible, not hidden.
Your simulated personas are great for pre-ship scoring, but the hesitation signal on a real face during a real task is a different class of data. I'd genuinely love to walk you through how the disagreement layer looks in a real study. Can we set up a call? → https://www.entropik.io/book-demo
@mridhu_varshini_ surfaced not resolved is the answer I was hoping for. weighting by how long a signal is sustained makes sense too, a six-second hold is a different animal from a flicker. no call promises mid-launch-week, but I'll poke at the demo materials. appreciate you writing this out properly.
Mira by Decode
@narek_keshishyan Glad it landed that way. "surfaced not resolved" is genuinely the design philosophy, not just a talking point.
No pressure on the call. When you do poke at the demo materials, if anything raises a question worth going deeper on, I'm here. Would be curious to hear what you think after you've had a look.
Congratulations on launch! A lot of the questions are about accuracy and privacy, so I'll ask a different one: how does the emotion reading hold up across cultures and languages? People show feelings differently depending on background, and my audience is fairly reserved by nature. Does the model account for that, or is it mostly calibrated to more expressive participants?
Mira by Decode
@alieksia Anastasiia, this is one of the most important questions anyone can ask about emotion AI — and one we've had to answer in practice across 150+ brands in markets where emotional expressiveness varies significantly.
A few things we do specifically for this:
Individual baseline calibration is not a universal norm. At the start of every session, Mira establishes a neutral baseline for that participant. What matters is deviation from their baseline, not deviation from an average. A reserved participant who shifts even slightly is flagged because for them, that shift is meaningful. You're not comparing them to a participant in a different country who naturally gestures more.
Multimodal cross-referencing helps here significantly. Voice tone and pacing convey emotional signals across cultures in ways that facial expressions alone don't. Someone from a more reserved cultural background who shows minimal facial movement will often convey more signals in their voice, hesitation, changes in pacing, or a drop in confidence. Mira reads both simultaneously.
We've trained specifically on diverse participant pools. Our models have been built on research across APAC, MENA, Europe, and the Americas, not just Western expressive populations. That said, we're transparent: some cultural nuances remain an active area for improvement, and we'd rather flag lower-confidence signals than assert false precision.
The output accounts for this contextually. Signals are always shown with confidence scores and duration markers. A researcher working with a reserved audience can set their own threshold for what's worth acting on.
If your audience is particularly reserved by nature, I'd love to show you a session with that kind of participant profile, specifically, it's actually where the multimodal approach shows its biggest advantage over facial-only tools.
Worth a call? https://www.entropik.io/book-demo
@mridhu_varshini_ thank you for taking the time to explain this so fully. The baseline calibration and the multimodal reading for reserved participants answer my concern well. I'll keep Mira in mind for when I'm at that research stage.
Mira by Decode
@alieksia , sure. We can reconnect anytime when the research stage comes in. Happy to be connected :)
Timbal AI
When Mira spots a say/feel mismatch and digs deeper on its own, how do you keep that follow-up from leading the participant? A moderator reacting to visible confusion can easily plant the doubt rather than uncover it. Is the probe neutral by design, or tuned per study?
Decode by Entropik
@david_vilalta David, this is the question we lost sleep over. A leading probe is almost worse than no probe — it plants the narrative.
A few design decisions we made:
Probes are behaviorally triggered but linguistically neutral. When Mira detects a signal — a hesitation, an emotional shift, a response that doesn't match the facial/voice pattern — it doesn't say "you seemed confused." It says something like "You paused there — tell me more about what was going through your mind." The trigger is emotional, the language is open.
No interpretive language in the probe. Mira never names the emotion it detected back to the participant. It asks outward, not inward. This is a deliberate constraint built into the probing engine.
Tunable per study type, not per participant. Concept testing probes are calibrated differently from usability or brand research — because the type of signal differs. But within a study, every participant gets the same probe structure. That keeps cross-participant comparisons valid.
Probe depth is capped. Max 2 levels of follow-up on any one signal — so the interview doesn't become an interrogation.
The honest caveat: no probe is perfectly neutral. But compared to human moderators who vary tone and word choice across 40 interviews — Mira is at least consistently neutral. Happy to show you a live session → https://www.entropik.io/book-demo
"Reads how people feel" is the interesting (and risky) part. When it detects hesitation mid-interview, does it adapt its questioning in the moment or just annotate for the researcher afterward? I'm running beta-user interviews right now and what I always miss is what people didn't say, I would love to know if you surface that.
Mira by Decode
@chielephant Anthony, you've named the exact thing most AI interview tools get wrong: they annotate after, when the moment is already gone.
Mira adapts in the moment. When it detects hesitation, an emotional shift, or a mismatch between what's said and how it's said, it probes during the conversation rather than as a post-hoc tag. "You paused there. Tell me more about what was going through your mind." The follow-up happens while the participant is still in the thought, not after they've moved on.
On what people didn't say: this is where the multimodal layer earns its place. Mira flags moments when vocal confidence dropped without a corresponding verbal signal, when attention shifted during a key question, and when facial affect changed but the spoken answer was flat. Those are the gaps, the things a transcript alone will never show you. They surface in the report as flagged moments with timestamps, so you can go directly to "here's where something was happening that wasn't being said."
For beta-user interviews specifically, that layer tends to be most valuable on product moments — the exact second someone's expression changes when they hit a friction point, even while they're still saying "yeah, this makes sense."
Would love to show you what this looks like on a real session, especially with beta interviews, the signal density is usually high → https://www.entropik.io/book-demo
How do you prevent Emotion AI from overinterpreting facial expressions or cultural differences in user reactions?
Mira by Decode
@robert_dimla We avoid cultural overfitting by rejecting static global priors and instead computing affect as temporal deltas against a per user baseline. Our attention layer gates visual action units against orthogonal acoustic and semantic vectors down-weighting isolated facial anomalies as noise unless validated across channels or via active verbal probing.
The Say-Do gap framing is sharp, and running recruit → moderate → analyze → report as one agent is the ambitious part. For a team that already has research infra, two setup questions: can I bring my own recruited participants into a Mira study, or is moderation locked to your 100M panel? And can I export the raw session data — transcripts plus the behavioral signals you capture — into our own repo, or does the analysis stay inside Mira's dashboards?
Mira by Decode
@noctis06 Good questions, both get asked by every team with existing infra, so let me be precise rather than salesy.
Bring your own participants: yes. Moderation isn't locked to our panel — you can pipe in your own recruited participants and Mira runs the same moderate → analyze pipeline on them. Two ways teams do this depending on setup:
Link-based: we generate a session link per participant, you distribute through your own recruiting/incentive flow, they land in the same moderated session experience.
API-based: if you've got a recruiting/panel system already, you can push participants in via API and we handle scheduling + session invites from there.
On export: yes, and I'll be specific about what "raw" means. CSV export gives you:
Full transcripts (turn-by-turn, timestamped)
The say-do gap annotations themselves (where the model flagged claimed vs. observed behavior divergence, with the reasoning trace)
Nothing stays trapped in the dashboard-only view. The dashboard is a convenience layer, not the source of truth — the source of truth is the same data you'd get in the export.