Why write about our own limits?
Every new research method gets oversold. We'd rather tell you where ours stops.
The scientific literature on synthetic research is remarkably clear on its weak spots. Here are the six that matter most, what the research says, and what we do about each one.
1. Lived experience and the body
The limit: language models can explain how people reason about a choice. They don't experience the physical and emotional friction of living it: the tired evening, the crying baby, the frustration of a form that won't submit. Researchers have warned that this missing embodiment limits what models can simulate (Harding et al., 2023).
What we do: we ground twins in real interviews, where that lived experience is described in people's own words. And for decisions that hinge on physical experience, we recommend confirming with real people.
2. Questions far from the data
The limit: research shows that digital twins perform well on topics related to what they were built from, and fall back on generic answers for topics with no link to their data. Toubia and colleagues (2025) call this the "information-free floor".
The nuance, and the opportunity: this doesn't mean twins can only repeat what's in the data. Echo is built for two things:
- Interpolation: going deeper on topics your data covers, making the data talk. Asking the follow-up questions nobody asked in the original study.
- Extrapolation: predicting reactions to new things, like a new price, a new contract or a new concept.
The key, confirmed by the research, is that new topics need to be adjacent to the data. You can ask an energy supplier's customers about their preferences for new contract types and pricing models: that's close to what they know and care about. You can't ask the same twins how they'd vote in a football club's board election and expect a meaningful answer.
What we do: we design studies within that adjacent space, and we're building predictive capabilities that extend it further.
3. Predicting one specific person
The limit: a synthetic sample can match the overall picture of a group while individual twins predict individual people poorly. Netzer and Sambandam (2026) show that reliability depends on the level of the decision: strong for group-level questions, weaker for one-to-one predictions.
What we do: we measure Echo's accuracy at study level, where teams actually decide: themes, segments, concepts and messages. We don't sell individual predictions.
4. Hyper-rational respondents
The limit: left to themselves, AI respondents can be too sensible. In usability research, synthetic users have sailed through confusing screens that real people got stuck on. Models tend to reason perfectly where real people get tired, distracted or irrational.
What we do: every Echo twin gets a personality. Some are impatient, some anxious, some sceptical, some expansive, some blunt. That shapes how they respond and keeps them from all behaving like the perfect, rational customer. We also include deliberate edge cases in every audience.
5. Telling you what you want to hear
The limit: language models can drift towards the answer a question seems to expect. Research on this "sycophancy" shows capable models bending conclusions to match the user's implied preference (Allen & Peterson, 2026).
What we do: our researchers design neutral, balanced questions with you, without hinting at the answer you hope for. And we validate studies against real interviews, not against expectations.
6. Garbage in, garbage out
The limit: twins reflect what they're built from. A biased or wishful audience description produces biased, wishful twins.
What we do: we help you write an honest brief. In one pilot, marketing, sales and insights wrote down only what they knew to be objectively true about their customers. That one-page brief reached 83% accuracy. Honest inputs matter more than big inputs.
So where does synthetic research shine?
With those limits in mind, here's where it adds the most value:
- Exploring early, before anyone has spent time or budget
- Screening many concepts and keeping only the strongest
- Sharpening questions before real interviews, so fieldwork confirms instead of explores
- Comparing segments side by side, in one study
- Reaching hard-to-reach groups: specialists, niche segments, markets you haven't entered yet
- Iterating the same day: new angle, new question, new run
- Making old research work again: turning past studies into an audience you can question anew
The rule of thumb
Use synthetic research to explore and narrow down. Use real customers to confirm the big bets. Echo is built for the first, and makes the second sharper.
The short version
Knowing the limits doesn't make synthetic research less useful. It's what makes it useful.
Frequently asked questions
What are the limitations of synthetic research? The main limits are missing lived experience, weaker answers on topics far from the underlying data, lower accuracy for predicting individuals, a tendency to be too rational, sensitivity to leading questions, and dependence on input quality.
Can digital twins predict new behaviour? Yes, when the new topic is adjacent to what the twins are built from, like new pricing for existing customers. Accuracy drops for topics with no link to the data.
When should I use real customers instead? For the biggest, highest-risk decisions, and for anything that depends on physical experience. Use synthetic research first to narrow down what to test.
Read the research behind this article: What the science says.
Try Echo free