Why does a language model need grounding?
A language model knows a lot about people in general. It knows nothing about your customers, until you show it.
Ask a model how a customer feels about switching energy supplier, and it answers from its own general picture of "a customer". Fluent, plausible, and generic. That's useful for brainstorming. It's not evidence about the people you actually sell to.
Grounding changes that. Instead of answering from memory, the model answers from real material.
What is RAG, in plain language?
RAG stands for retrieval-augmented generation. It's a two-step process:
- Retrieve. Before answering, the system searches a body of documents for the passages most relevant to the question.
- Generate. The model answers with those passages in front of it.
Think of a well-briefed actor. Before every line, they reread the real interviews with the character they're playing: what that person worries about, how they talk, what they'd never say. The performance stays true to the character because the source material is always at hand.
Where does Echo look for context?
Every Echo twin draws on three layers of data:
| Layer | What it contains | What it adds |
|---|---|---|
| Your data and reports | Interview transcripts, survey answers, segmentations, past studies | Your customers' own words, concerns and contradictions |
| Public data | Published research, market statistics, public sources about the audience | Context about the wider market and the people in it |
| Echo's research data | Our own research knowledge, built from our own studies | Proven patterns that make every new audience sharper |
Your data is used only for your studies. It's never used to train AI models, and never shared with anyone else.
How does Echo find the right passages?
At a high level:
- Your documents are split into passages and indexed by meaning, not by keywords. A question about "monthly costs" finds passages about "what I pay each month", even when the words differ.
- Each question pulls the passages that matter for that topic. A question about pricing retrieves what customers said about price, not everything they ever said.
- Retrieval is personal to each twin. This is the part that makes Echo different, so it gets its own section.
Why does every twin get different context?
Because real customers aren't one voice.
Echo discovers the sub-groups in your data: the groups that emerge from what people actually say. Each twin belongs to one of them and has its own background and personality. When a twin answers, it draws on the passages from voices like its own.
The anxious first-time buyer draws on what anxious first-time buyers said. The experienced, price-focused customer draws on what experienced, price-focused customers said. Same question, different context, different answers.
That's what makes Echo's output rich and diverse. Instead of one polished average, you get the spread of opinions, doubts and trade-offs you'd find in a real group of customers. Segments sound different from each other, because in your data, they are.
What grounding is, and what it isn't
It is: - Answers anchored in real language and real concerns - Less generic filler, fewer invented details - Twins that stay true to their segment across a whole study
It isn't: - Training a model on your data. Your data stays in your project and never trains AI models. - A quote-per-answer citation system. Twins speak in their own words, informed by the sources. - A guarantee that every passage is used. Retrieval picks the most relevant passages for each question.
How do we know it works?
Because we measure it. Grounded twins are scored against real human interviews with the Echo Accuracy Score, which reaches 85–93% across our customer engagements. The scientific research points the same way: twins built from rich, real data clearly outperform twins built from demographics alone (see What the science says).
The short version
A model on its own guesses. A grounded twin knows where to look. That's the difference between a plausible answer and an accurate one.
Frequently asked questions
What is retrieval-augmented generation (RAG)? RAG is a method where an AI model first retrieves relevant passages from a set of documents and then answers with those passages as context, instead of answering from memory alone.
What data does Echo use to build digital twins? Echo combines three layers: your own data and reports, publicly available data and Echo's own research data. Your data is used only for your studies.
Is my data used to train AI models? No. Your data is used only to build your audiences and run your studies. It's never used to train AI models, Google's or Echo's own.
See it on your own data. Try Echo free or read what data you need to start.
Try Echo free