← Research notes
Validation · ~6 min

Measuring the echo: how we score synthetic answers against real ones

By the Echo research team · Updated 1 October 2026
HOW THE ECHO ACCURACY SCORE IS CALIBRATEDReal group A vs real group BHuman-to-human agreement: the natural ceiling = 100%Echo vs real groupEcho agreement, relative to that ceiling = 85–93%range across runsThemesFeelingMeaningRangeConsistencyCompared at study level, averaged over many parallel simulation runs

Every team that looks at synthetic research asks the same first question: how do you know it's right?

It's the right question. A synthetic interview that sounds convincing is easy to produce. Any language model can write a fluent paragraph in the voice of a 34-year-old parent. Whether that paragraph says what real parents would say is a different matter. Fluency is cheap. Evidence is not.

In A Brief History of Time, Stephen Hawking cites the philosopher of science Karl Popper:

"A good theory is characterized by the fact that it makes a number of predictions that could in principle be disproved or falsified by observation."

We hold Echo to the same standard. Every Echo study makes predictions about your customers: what they'll raise, how they'll feel, where segments will differ. And every one of those predictions can be checked against real people.

So we built Echo around a simple rule: we don't ask you to trust the output. We measure it.

Start with a real study

The Echo Accuracy Score starts where all good validation starts: with real people.

We take a study that was also run with human respondents: the same audience, the same questions. Then we run it in Echo and compare the two sets of interviews side by side. Not one answer against one answer, but the whole study against the whole study: what came up, how people felt about it, and whether the picture you'd walk away with is the same.

That last part matters most. Nobody makes a decision based on one quote. Teams act on the patterns: the themes that keep coming back, the frustrations that cluster in one segment, the idea that half the room loves and half the room hates. That's what we score.

What we compare

The score looks at a study from several angles at once:

  • Themes. Do the same topics surface on both sides? If real customers keep bringing up delivery times, do Echo's twins as well?
  • Feeling. It's not enough to mention the right topics. If real customers are frustrated about price, Echo's twins should be frustrated too, not politely neutral.
  • Meaning. Do the answers say the same things, even when they use different words?
  • Range. Real groups are messy. Some people answer in a sentence, some in a paragraph. Some love a concept, some hate it. We check that Echo keeps that variety instead of collapsing into one polished average answer.
  • Consistency. Does each synthetic respondent stay in character from the first question to the last?

Together, these tell you whether the conclusions you'd draw from Echo match the ones you'd draw from real interviews.

The part most benchmarks skip: real people disagree too

Here's the problem with a simple "match percentage". Ask two groups of real customers the same questions, and they won't give identical answers either. People are inconsistent, contradictory and wonderfully varied. A synthetic study can never match one human group perfectly, because another human group wouldn't match it perfectly either.

So we don't score Echo against perfection. We score it against how much real people naturally differ from each other on the same study. First we measure that human-to-human baseline. Then we express Echo's agreement relative to it.

Think of grading a translation. You wouldn't grade it against one "perfect" text. You'd compare it to how much two expert translators naturally differ from each other. If the machine is as close to the experts as the experts are to each other, that's the bar. Echo is scored the same way.

This matters because it makes the number honest. A score that ignores human variation either flatters the model or punishes it unfairly. A score calibrated against real disagreement tells you what you actually want to know: is this as close to my customers as another group of my customers would be?

How to read 85–93%

Across our customer engagements, Echo reaches an accuracy of 85–93% on well-set-up studies. Three things are worth knowing about that number.

It's a study-level score. It tells you how close the overall picture is: the themes, the sentiment, the conclusions. It is not a promise that every individual answer is right. An individual twin can say something off even when the study as a whole lands.

It's an average of many runs. Language models are not deterministic: the same study never plays out exactly the same way twice. So we don't score a single run. We run many accuracy simulations in parallel and report the average across all of them. The spread between runs is small, which is why we report a range rather than one exact figure.

It depends on the setup. A sharp question and a well-defined audience score higher than a vague one. That's true for human research too.

And the other 7–15%?

We don't hide the gap. We show it.

Every validated study comes with a view of what Echo matched, what it missed, and what it added compared to the real interviews. That gap is often the most interesting part. It's where culture, emotion and plain human irrationality show up: the behaviour no dataset predicted. More on that in Is 100% accuracy the goal?.

Why this matters for your decision

The point of the Echo Accuracy Score isn't a nice number for a slide. It's the confidence to act.

When you know how closely Echo tracks real customers, and where it doesn't, you can decide what to do with the result. Use it to kill weak concepts early. Use it to sharpen the questions you take into real interviews. Use it to move an internal debate from we think to we tested.

Real customers stay the ground truth. Echo is how you get to them with better questions, faster.

Next: What data do you need to build accurate digital twins? · What the science says

Want to see the score on your own audience? Every plan starts with real customer interviews, so you see your own Echo Accuracy Score. Explore my options

Try Echo free