Synthetic research can sound like a leap of faith. Asking an AI how your customers feel, and trusting the answer?
It isn't a leap. Over the past few years, researchers at Stanford, Harvard, Wharton, Columbia, MIT and Brigham Young University have put the idea to the test, again and again, against real human data. The results are strong enough that the field now has its own vocabulary: silicon samples, Homo silicus, digital twins. The same research is equally clear about where the limits lie.
We read this literature closely and keep following it. It shapes how we build Echo. But we don't just follow it: we also have our own view on how this technology should work for teams making real decisions, and that view is what you'll find in the product.
Here is what the research says, in plain language.
1. Language models can stand in for specific groups of people
Argyle et al., "Out of One, Many: Using Language Models to Simulate Human Samples", Political Analysis (2023).
The researchers gave a language model the detailed backgrounds of real survey respondents (age, gender, political views, region) and asked it to answer as them. The simulated answers reproduced how those real groups answered, including the relationships between demographics, opinions and choices. The authors call this algorithmic fidelity, and the simulated groups silicon samples.
Why it matters: the core idea behind digital twins. Who the twin is changes what it says, in the same direction real people differ.
2. Simulated people share our human instincts
Horton, "Large Language Models as Simulated Economic Agents: What Can We Learn from Homo Silicus?" (2023).
Horton ran classic behavioural economics experiments on simulated participants. A famous one: a hardware store raises the price of snow shovels the morning after a blizzard. Just like real people, the simulated agents mostly judged it unfair. Give them different political profiles, and their judgement shifted the way real people's does.
Why it matters: synthetic respondents don't just produce plausible sentences. They carry the moral intuitions and mental shortcuts that drive real choices.
3. Language models can do market research
Brand, Israeli & Ngwe, "Using LLMs for Market Research" (Harvard Business School, 2023).
The authors simulated more than 10,000 purchase decisions for everyday products like toothpaste, in about half an hour, for a few dollars. The results behaved like real markets: demand dropped as prices rose, willingness to pay for features matched established benchmarks, and brand preferences showed up as real price premiums.
Why it matters: this is the closest academic test of what Echo does every day: putting concepts, prices and propositions in front of synthetic customers.
4. Simulations can predict the results of new experiments
Hewitt, Ashokkumar, Ghezae & Willer, "Predicting Results of Social Science Experiments Using Large Language Models" (Stanford, 2024).
The team asked a model to predict the outcomes of dozens of social science experiments, including studies that were never published, so the model couldn't have seen the answers. Its predictions tracked the real outcomes closely, and matched or beat human experts.
Why it matters: synthetic research isn't only replaying what's already known. It can anticipate how people react to something new, which is exactly what concept and message testing needs.
5. Twins built from real interviews come remarkably close to real people
Park et al., "Generative Agent Simulations of 1,000 People" (Stanford & Google DeepMind, 2024). Toubia et al., "Twin-2K-500: A dataset for building digital twins of over 2,000 people" (Columbia, 2025).
Park and colleagues interviewed over 1,000 real people in depth and built an AI agent for each one. The agents answered a major social survey about as consistently with the real person as that person did when re-taking the survey two weeks later. Agents built from rich interviews clearly outperformed agents built from demographics alone. Toubia and colleagues confirmed the pattern on more than 2,000 people: twins do well on topics related to what they were built from, and fall back on generic answers when they have nothing to go on.
Why it matters: two points at the heart of Echo. Grounding in real qualitative data is what makes twins accurate. And the right benchmark is how consistent real people are with themselves, which is why we score Echo against human-to-human variation.
6. AI can widen the funnel of good ideas
Girotra, Meincke, Terwiesch & Ulrich, "Ideas are Dimes a Dozen: Large Language Models for Idea Generation in Innovation" (Wharton, 2023).
In a head-to-head tournament, AI-generated product ideas were rated higher on purchase intent than ideas from elite design students, and were far more likely to land among the very best. Human designers kept an edge on truly radical ideas.
Why it matters: in innovation, what counts is the quality of the best idea, not the average. Synthetic research lets teams test many more options before choosing, and keep humans where originality matters most.
The other side: what the research warns about, and what we do about it
Good science also shows where things break. Four findings shape how Echo is built.
Default models lean towards some groups. Santurkar et al. (ICML, 2023) showed that off-the-shelf models over-represent the views of highly educated, higher-income groups. What we do: Echo never relies on a model's default "average person". Every twin is built from your audience description or your own customer data.
Averages can look right while variety disappears. Dillion et al. (Trends in Cognitive Sciences, 2023) and Bisbee et al. (Political Analysis, 2024) found that synthetic answers can match the average while missing the spread of real opinions. What we do: we preserve variation on purpose: distinct personalities, sub-groups discovered from your data, and deliberate edge cases. And our accuracy score checks the range of answers, not just the average.
Aggregate accuracy doesn't guarantee individual accuracy. Netzer & Sambandam (2026) showed that a synthetic sample can match the overall picture while individual twins predict poorly, and that reliability depends on the type of decision. What we do: we're explicit that Echo's accuracy is measured at study level, the level where teams make decisions: themes, segments, concepts. We don't sell individual predictions.
Models can tell you what you want to hear. Research on sycophancy (Allen & Peterson, 2026) shows models can bend their conclusions towards the answer a user seems to expect. What we do: our researchers design neutral, balanced questions with you, and we validate studies against real interviews rather than against expectations.
The short version
The research doesn't say AI can replace your customers. It says something more useful: when synthetic respondents are grounded in real data, kept diverse and measured against real people, they reproduce human attitudes and behaviour closely enough to explore, test and decide much faster. That is what we built Echo to do. And as the field moves, we keep reading, testing and building on it.
Further reading: other scientific papers on synthetic research and digital twins
Foundations - Argyle, L. P., Busby, E. C., Fulda, N., Gubler, J. R., Rytting, C., & Wingate, D. (2023). Out of One, Many: Using Language Models to Simulate Human Samples. Political Analysis, 31(3). - Horton, J. J. (2023). Large Language Models as Simulated Economic Agents: What Can We Learn from Homo Silicus? NBER Working Paper. - Aher, G., Arriaga, R. I., & Kalai, A. T. (2023). Using Large Language Models to Simulate Multiple Humans and Replicate Human Subject Studies. Proceedings of ICML. - Park, J. S., O'Brien, J. C., Cai, C. J., Morris, M. R., Liang, P., & Bernstein, M. S. (2023). Generative Agents: Interactive Simulacra of Human Behavior. Proceedings of ACM UIST.
Digital twins of individuals - Park, J. S., Zou, C. Q., Shaw, A., Hill, B. M., Cai, C., Morris, M. R., Willer, R., Liang, P., & Bernstein, M. S. (2024). Generative Agent Simulations of 1,000 People. arXiv:2411.10109. - Toubia, O., et al. (2025). Twin-2K-500: A Dataset for Building Digital Twins of over 2,000 People Based on Their Answers to over 500 Questions. arXiv:2505.17479.
Market research and consumer behaviour - Brand, J., Israeli, A., & Ngwe, D. (2023). Using LLMs for Market Research. Harvard Business School Working Paper 23-062. - Netzer, O., & Sambandam, R. (2026). Synthetic Data in Marketing Research: How to Evaluate and When to Use It. arXiv:2609.13995.
Prediction and experiments - Hewitt, L., Ashokkumar, A., Ghezae, I., & Willer, R. (2024). Predicting Results of Social Science Experiments Using Large Language Models. Working paper, Stanford University.
Innovation and strategy - Girotra, K., Meincke, L., Terwiesch, C., & Ulrich, K. T. (2023). Ideas are Dimes a Dozen: Large Language Models for Idea Generation in Innovation. Wharton working paper. - Csaszar, F. A., Peterson, A., & Wilde, D. (2026). The Strategic Foresight of LLMs: Evidence from a Fully Prospective Venture Tournament. arXiv:2602.01684.
User experience research - Xiang, W., et al. (2024). SimUser: Simulating User Behavior with LLMs for Mobile App Usability Evaluation. Proceedings of ACM CHI.
Limits and critiques - Santurkar, S., Durmus, E., Ladhak, F., Lee, C., Liang, P., & Hashimoto, T. (2023). Whose Opinions Do Language Models Reflect? Proceedings of ICML. - Dillion, D., Tandon, N., Gu, Y., & Gray, K. (2023). Can AI Language Models Replace Human Participants? Trends in Cognitive Sciences, 27(7). - Harding, J., D'Alessandro, W., Laskowski, N. G., & Long, R. (2023). AI Language Models Cannot Replace Human Research Participants. AI & Society. - Bisbee, J., Clinton, J. D., Dorff, C., Kenkel, B., & Larson, J. M. (2024). Synthetic Replacements for Human Survey Data? The Perils of Large Language Models. Political Analysis. - Allen, & Peterson, A. (2026). Intelligence Without Integrity: Why Capable LLMs May Undermine Strategic Decision-Making. arXiv:2602.20440.