A synthetic respondent is an AI model standing in for a human survey participant. Rather than recruiting people and collecting their answers, you prompt a large language model to simulate the responses of a described audience. The idea has spread quickly because it promises what every researcher wants and rarely gets: instant, almost-free "data". The trick is knowing which of those quotation marks to take seriously.
Where synthetic respondents earn their place
The honest, valuable uses are about preparation, not conclusions. The strongest is pre-testing. Before a survey goes to real people, you can have AI run through it and flag the things that quietly ruin studies: leading or double-barrelled questions, missing answer options, confusing logic, dead ends in the flow. Catching these before fieldwork protects your real responses from being wasted on a broken instrument. Synthetic runs are also useful for generating hypotheses to test properly later, for rehearsing an analysis plan so your spreadsheet is ready before the data lands, and for rough, clearly-labelled illustrations when no real data exists yet.
Where they quietly mislead
The fundamental problem is that a language model has no opinions, experiences, or stakes. It produces the most statistically plausible text, which is not the same as what your customers actually think. That gap creates several specific traps. It regresses to the mean of its training data, so niche, local, or recent realities get flattened into a generic internet-average view. It carries the biases of that data, often under-representing exactly the groups you most need to hear. It is sycophantic, tending to agree with the framing of your prompt, so a leading question yields a leading answer with no friction. And most dangerously, it is confident: it will hand you a clean, decisive "62% prefer option A" that looks like a finding but is closer to a hallucination dressed as data.
The line you should not cross
Synthetic respondents can rehearse a study; they cannot be the study. The moment you present simulated answers as evidence of what real people believe — to a client, to leadership, in a decision that moves money — you have crossed from a useful tool into a fabricated result. Real research derives its authority from contact with reality: actual people, with actual experiences, answering on their own behalf. Remove that, and you have a well-written guess. Used well, the workflow is a loop: rehearse with synthetic respondents to sharpen the instrument, then field the real survey and collect answers from genuine participants — and let those real answers, not the simulated ones, drive the decision.
Common Misconceptions
Most people think
"Synthetic respondents can replace real surveys and save the whole cost of
fieldwork."
Actually
They replace the cost of a pilot, not the value of real data. Simulated
answers can make a study better, but they carry no information about what
your actual customers think — only what the model predicts is plausible.
Most people think
"If the AI is trained on huge amounts of human data, its answers are
basically representative."
Actually
Training data is not a representative sample of your audience. It
over-weights what is common online and under-weights the specific,
local, and recent — the very things most research is trying to measure.
Common mistakes
The biggest mistake is laundering a guess into a number: running a synthetic panel, getting a tidy percentage, and presenting it as if real people said it. The second is using synthetic respondents on exactly the questions they are worst at — novel products, local markets, emotional or sensitive topics — where there is no reliable pattern to draw on and the model simply invents a confident answer. The third, less obvious, is skipping the real study entirely because the synthetic one was so easy, and never noticing that you stopped doing research and started consulting a mirror. Treated as a rehearsal tool, synthetic respondents make real research faster and sharper. Treated as a shortcut around real people, they quietly replace evidence with fiction.