A simulation product is only as trustworthy as the evaluation behind it. But how does your simulation product give you the context you need on that evaluation?
Predictive Answers, our simulation product, answers research questions by simulating the segments in your audience. Predictive Answers uses your own research sources – interviews, surveys, support tickets, CRM data – and attaches citations and a confidence score to every answer, giving you full transparency that other simulation tools don't offer (more on that later).
Behind the scenes, we pressure-test the simulation from Predictive Answers at increasingly demanding levels. One of the most challenging tests is how well we can predict an individual person, idiosyncrasies included. If simulation holds up there, all of the segments above it gets easier. So we ran a study at that individual level to see how we held up across different question types – on questions types that the twins are best suited to answer and extended to use cases beyond those.
On our consumer-signal framework, our framework of the questions that matter most for consumer research, the twins hit 82%. Pushed beyond the framework into hard-to-predict territory, they still scored as high as 76%.
The gradient from this study maps directly onto how the confidence scoring in Predictive Answers works. Confidence tracks evidence strength, and this study illustrates why it must.
Here's the study and what it taught us about how to evaluate a simulated answer.
The Experiment: A Two-Instrument Simulation Protocol
Instrument 1: The grounding interview
We built digital twins of 10 real people, each grounded in nothing but a 45-minute AI voice interview on Persona. Notably, this is half the length of the 2-hour interviews used in Generative Agent Simulations of 1,000 People, the landmark Stanford study on generative agents. One of our questions was whether interview quality and instrument design could substitute for raw interview length.
Each of our 10 participants completed a 45–60 minute AI voice interview run on Persona, conducted entirely by our AI interviewer. The interview guide is our own adaptation of established life-history methodology, extended with consumer sections academic protocols don't cover. In one conversation, it covers life story and turning points, relationships and daily life, political and religious worldview, health and coping, money and work, future aspirations and values, and everyday consumer habits – from the brands they buy to what drives their purchase decisions.
The AI interviewer asks adaptive follow-up questions based on what participants reveal, which means every transcript captures not just facts but decision logic – how each person actually reasons about money, risk, and values.
Instrument 2: The held-out evaluation battery
Later, the same participants completed a decision survey the twin never saw during construction – spanning identity, personality self-ratings, and behavioral scenarios: surprise windfalls, guaranteed-money-versus-gamble choices, investment allocations, job offers, privacy trade-offs, and civic attitudes.
Every number below is out-of-sample: the twins were never tuned against the people or the questions being scored.
The Consumer-Signal Framework: 82%
At the center of the evaluation is a framework we developed for the question types consumer research asks most – the decisions that drive purchasing, loyalty, and values-based behavior. Across ten questions, it covers how people spend and manage money, what they're willing to pay for, how they handle financial risk, and the values and worldview that shape those choices.
82%
Twin accuracy on the 10-question consumer-signal framework, across 10 participants
One twin scored a perfect 10 out of 10 on this set. Three more scored 9 out of 10.
This is the territory that matters if you're using simulation for consumer insight – and it's also, not coincidentally, the territory a life interview is best equipped to reveal. People talk about their money habits and values constantly, in stories, complaints, and plans. The framework works because it targets attitudes people have rehearsed their whole lives.
Beyond the Framework: How Twins Handle Questions They Weren't Built For
The framework tells you where simulation is strongest. But we also wanted to know how far the twins could stretch – so we deliberately extended the battery into two categories of questions that are harder to predict, or only loosely related to what a life interview reveals.
Category 1: Forward-looking attitudes. How participants feel about AI in daily life, whether it will reshape their careers, whether universities should allow it in coursework, their general risk preference, and where they sit on a five-point political scale. These are real attitudes, but they're constructed attitudes – people are still forming them, and they shift with framing and mood. Even with these added, the twins held at 76% across the expanded 15-question set.
Category 2: Fine gradations. Questions where the answer is a matter of degree: how much confidence in the press – some, or very little? Do you worry somewhat or strongly? Stable corporate job, or negotiate first? When the twin misses these, it's almost always by one adjacent notch – predicting "some confidence" when the person said "very little." The direction is right; the calibration is off by one step. With both hard categories included, accuracy across all 21 questions was still 70%.
82% → 76% → 70%
Framework accuracy, then with forward-looking attitudes added, then with fine gradations added
For a twin built from a single 45-minute conversation, holding 70%+ while answering questions it was never designed for is the more surprising result. It suggests the interview is doing real work: capturing not just facts a person states outright, but the reasoning patterns that let a twin extend to territory the conversation never directly covered.
Why the Gradient Exists
This pattern echoes a much older finding in survey research. Political scientists John Zaller and Stanley Feldman's "A Simple Theory of the Survey Response" argues that people don't walk around with fixed, pre-packaged answers stored in their heads. Instead, they construct an answer on the spot, out of whatever considerations happen to be accessible at that exact moment – a recent experience, a mood, whatever framing the question happened to use.
This is one cause of a phenomenon that researchers are familiar with: even the same participant can give different responses to the exact same question.
If that's true for the real person, it shapes what simulation can do: the more stable and rehearsed the underlying attitude, the more predictable the answer. Money habits and values are rehearsed daily. A stance on self-driving cars might be constructed for the first time when the question is asked – by the participant and the twin alike.
Key Insight
The Hardest Test: Open-Ended Voice Responses
If structured survey prediction is already difficult, open-ended voice responses are harder by an order of magnitude. There's no constrained answer space, no option list to anchor to, and the evaluation itself is inherently subjective – we used an LLM judge to score each twin response against the real participant answer on stance, content, and style, which captures meaning rather than exact wording, but introduces its own uncertainty. These scores should be read as directional, not definitive.
We believe open-ended prediction is among the hardest challenges in digital twin research with a limited dataset. A twin needs to not just hold the right opinion, but express it in the right voice, at the right level of specificity, touching the right details – all from a single 45-minute interview. That's a high bar.
Which is why the results surprised us. When we asked the same participants free-form questions in a follow-up interview – about the government's role in the environment, how they feel about the media, their career values, and their spending habits – the twins scored 75–85% on those topics consistently across participants. The same question types that scored highest on the survey scored highest in free conversation too. The gradient isn't an artifact of the multiple-choice format. It reflects something real about how stable certain attitudes are, and how reliably a good interview surfaces them.
Why Confidence Scores Work in Predictive Answers
Most research questions aren't about one person – they're about audiences. Predictive Answers simulates those audiences from your own first-party research – interviews, surveys, CRM, support tickets – letting you ask any question and get a cited, confidence-scored answer in seconds.
At the segment level, the idiosyncratic misses we measured on individual twins wash out while shared patterns compound, which is why 82% on individual twins is the floor, not the ceiling, for what Predictive Answers delivers.
The gradient from this study maps directly onto how its confidence scoring works:
- Confidence tracks evidence strength, and this study shows why it must. Our results map exactly where simulation is reliable: Consumer-Signal Framework answers predicted at 82%; constructed attitudes at 76%; fine gradations at 70%; one-off specifics lower.
- When Predictive Answers scores an answer, it's measuring the same thing – how rich, direct, and consistent the supporting evidence is. A question about spending psychology backed by twenty interview passages scores high. A question the corpus barely touches scores low, visibly, instead of hiding behind a fluent answer.
- Contradictions lower confidence – because they should. Our weakest twins came from interviews where self-description and behavior pointed in different directions. Predictive Answers treats contradictory evidence the same way: it penalizes the score rather than picking a side silently.
- Grounding beats generic. In our calibration work, prompts grounded in interview evidence consistently outperformed stripped-down ones. Generic synthetic respondents – built from LLM training data and internet averages instead of your customers – lack that grounding entirely. Predictive Answers grounds every segment model in first-party research: your interviews, surveys, support tickets, and CRM data.
- Every claim cites its evidence. In our experiment, each twin prediction cited the specific interview passage behind it. Predictive Answers works the same way: every simulated answer links to the exact interview clips and survey responses supporting it, so you can check the reasoning yourself.
This is a small-scale example – 10 participants, one demographic segment – and the results should be read as directional, not as a definitive or scientifically validated benchmark. Even so, we think they offer a useful signal for any brand or research team weighing synthetic research and digital twins against the alternative of not testing at all.
See simulation grounded in your own research
Predictive Answers builds cited, confidence-scored segment models from your interviews, surveys, and CRM data – so you know what your customers would say before you ask them.
Book a demo