Updated 24 September 2026: This article has been revised to bring in a wave of academic research published between January and September 2026 that directly tests whether AI-generated "digital twins" can substitute for human respondents. The results are consistent across independent research teams and datasets: synthetic respondents fail at individual-level prediction and at segment-level decisions, the exact tasks most commercial research depends on. The new evidence is in the section below.
Synthetic respondents, AI personas built to answer surveys and interviews as if they were real people, hold up for a narrow band of low-stakes tasks: piloting a questionnaire, or checking that a landing page covers the categories every competitor already covers. For anything that requires predicting how a specific person thinks, or which customer segment actually differs from another, the 2026 academic record says they do not. A cross-domain benchmark against the U.S. General Social Survey and the World Values Survey found that no large language model beat a simple demographic lookup table at predicting individual responses, and that segment-targeting decisions built on synthetic data picked the wrong segment in roughly half of U.S. cases tested.
The pitch is still everywhere. Feed demographic criteria and research questions into a platform, and receive detailed responses from synthetic personas that behave like your target customers, in an afternoon instead of weeks. Synthetic respondents were flagged as one of the major trends market researchers were told to watch in 2026, and companies like Evidenza and NIQ BASES have built platforms around exactly this promise.
For academics, synthetic respondents remain close to a no-go zone. For market researchers, the debate has moved on from "is this innovation or a shortcut" to "which narrow cases does this actually hold up for, and which ones is the evidence now closing off." This article goes through what the newest research shows, and where the technology still earns its place.
What are synthetic respondents and why the sudden interest?
Synthetic respondents are AI-generated personas that answer research questions as if they were real people. You provide the system with demographic characteristics, psychographic profiles, and behavioural patterns drawn from existing data, and the AI simulates how people matching those criteria would respond to surveys, interviews, or concept tests. The technology builds on large language models trained on vast amounts of human-generated text: ask a synthetic respondent representing a 35-year-old mother of two from Manchester what she thinks about a new grocery delivery service, and the model draws on patterns from millions of similar conversations to generate a plausible-sounding answer.
The market research industry paid serious attention going into 2026, with Rival Group, a leader in AI-accelerated conversational research, naming synthetic respondents a major trend. Companies from Fortune 500 enterprises to startups experimented with the technology for everything from concept testing to brand perception studies.
The appeal is obvious. Recruiting 20 qualified participants for user interviews can take weeks and cost thousands. Running a survey with 500 respondents requires panel access, incentives, and data cleaning. Synthetic respondents promise to compress timelines from weeks to hours and cut costs by 90% or more, which is exactly the pricing pressure pushing agencies toward the commoditisation trap if they compete on the same terms rather than on defensible depth.
What does the 2026 academic research actually show?
Until recently, most of the pushback on synthetic respondents came from methodologists' general unease and a handful of anecdotes. That changed in 2026. At least three independent research teams, working with different datasets and different model families, published quantified tests of exactly how well AI-generated respondents predict real human answers. The numbers converge on the same conclusion: synthetic respondents can approximate a population's average opinion in some conditions, but individual-level prediction and segment-level decisions, the two things most commercial research is actually used for, remain unreliable.
"When Synthetic Users Fail" (Chen, Zhu and Zheng, Stevens Institute of Technology and University of Massachusetts Boston, arXiv preprint, July 2026) is the most direct test to date. The researchers benchmarked four large language models against the U.S. General Social Survey and the 63-country World Values Survey, scoring each model against a simple non-LLM baseline built from held-out demographic data. Two failures showed up consistently across every model and dataset tested:
- No model beat the baseline. At predicting individual-level responses, "no LLM beats even the strongest baseline," and on cross-cultural values every model tested was 11 to 22 percentage points less accurate than a plain demographic lookup table.
- Demographic over-determination. Political affiliation explains roughly 1.5% of the real variation in people's confidence in banks, but up to 67% of the variation the models generated. The models treat demographic identity as a far stronger predictor of attitudes than it actually is, distorting nearly every question-group combination tested.
- Segment-targeting errors that would misdirect a real decision. On tasks that mimic how a brand team picks which customer segment to target, the models inflated the gap between segments two to fourfold, and would have led a team to the wrong segment in around 50% of U.S. cases and most cross-cultural scenarios.
"Leaving Insight to Digital Twins?" (Kaiser, Kaiser, Schallner, Manewitsch and Rau, published in NIM Marketing Intelligence Review, 2026) tested synthetic respondents against real consumer survey data across a brand-selection and attitude-rating scenario. Synthetic answers matched real brand choices only 79% of the time, and on 7-point rating scales the synthetic responses deviated from real ones by an average of 1.2 points, consistently skewed more positive. The synthetic responses were also significantly less variable than the real ones, especially for well-known brands. The authors' own conclusion: the technology is "most suitable for early-stage concept testing or lower-stakes applications," not precision-dependent decisions.
"Synthetic Personalities" (Kinzinger and Hartmann, TUM School of Management, Technical University of Munich, arXiv preprint, 2026) ran the most generous possible test: 500 real participants drawn from a 16,055-person panel, fed as much individual data as the researchers could give the model (full response histories, not just demographics), scored across 2.1 million synthetic responses. Even in this best case, the highest correlation between a synthetic respondent's answers and the real person's actual answers topped out at r = 0.59, a moderate relationship, not one you would want to underwrite a launch decision with. Giving the model a person's entire prior survey history over bare demographics lifted accuracy by only 5.2 percentage points on average.
These findings sit on top of earlier warning signs. A 2025 study in Political Analysis comparing synthetic ChatGPT responses to real human survey data found that 48% of coefficients estimated from AI responses were statistically significantly different from their human counterparts, and among those cases, the sign of the effect flipped 32% of the time. The 2026 research does not overturn that finding. It quantifies exactly where and how badly it holds, across bigger benchmarks, more model families, and a wider range of real-world decision tasks.
Industry doubts: vendors, researchers and ethics
The academic evidence lands on top of concerns the industry itself had already raised. Research World warned that synthetic data can lead to misleading conclusions if poorly generated or applied to the wrong contexts, noting the data tends to be too uniform and clean compared to real human responses.
Merrill Research documented a telling example. They created synthetic design engineers and asked them about sustainability in microprocessor vendor selection. The synthetic respondents gave textbook answers about the importance of sustainability. The real engineers said: "Sustainability matters, but not when we can't get the parts we need for months on end." This critical insight about prioritising availability over sustainability during supply chain disruptions was completely invisible to the AI.
Ethical concerns are equally significant. Without transparency about the use of synthetic data, clients may not realise their insights come from AI rather than real people. This raises fundamental questions about research integrity and trust. Industry bodies like the Market Research Society are working on guidelines, but regulation lags behind the technology.
Perhaps most telling is sentiment from researchers themselves. Rival Group's own study found that 42.75% of market researchers are "not excited" about using synthetic respondents, despite enthusiasm for other AI applications in research. That scepticism now has a growing body of quantified evidence behind it.
Our view: understanding the fundamental limitations
We have experimented with synthetic respondents internally at Skimle, for example to create dummy data for analysis (e.g., consultation responses to a new mall building project, and a fictional Due Diligence on ToyMaker producing toys for Santa Claus), and to get e.g., Claude Code to assess our website and give feedback on what do add, change or delete. Through this work and through discussions with market research companies and companies buying market research, we have started to develope a perspective on where the technology can add value and where it fails fundamentally. The 2026 research summarised above matches what we have seen hands-on: synthetic respondents are reasonable at telling you whether something matches an established pattern, and unreliable the moment a question depends on an individual person's history, taste, or a genuinely new situation.
Limitation 1: Trained on the past, blind to genuinely novel experiences
Synthetic respondents are fundamentally backward-looking. They are trained on existing data about how people have reacted to existing products, features, and experiences. This makes them reasonably good at telling you whether something matches established patterns, but terrible at evaluating genuinely novel concepts.
We use synthetic respondents internally to validate non-differentiating elements. Does our landing page include all the critical components that high-performing SaaS websites typically have? Are our blog posts on thematic analysis covering the standard topics that researchers expect? For these "solved problems" where we are trying to emulate existing best practices, synthetic feedback can be useful.
But imagine you have developed a genuinely innovative feature that changes how people think about qualitative data analysis. Something that triggers a "wow, I never thought about it that way" response in real users. Synthetic respondents will not spot this. They lack the underlying human experiences and cognitive processes that make unexpected innovation resonate. They evaluate new things through the lens of old patterns.
The same applies to research seeking genuinely novel insights. If you are exploring an emerging market, understanding evolving customer needs, or investigating how people adapt to new technologies, synthetic respondents will give you yesterday's wisdom, not tomorrow's understanding. Real qualitative research exists precisely to discover what you do not already know. Synthetic respondents can only reflect what the training data already contains in some shape or form.
Limitation 2: Credible responses without lived experience
When we tested synthetic respondents on questions about analysing open text survey responses, they generated articulate, plausible answers. They correctly identified pain points like "time-consuming manual coding" and "difficulty identifying meaningful patterns." These are real problems that appear frequently in research about research methods.
But here is what they missed: the specific frustration of being 200 responses into coding and realising your category framework is fundamentally flawed and you need to start over. The moment of doubt when a quote seems to fit two categories equally well. The satisfaction of spotting an unexpected pattern that no one in your team anticipated. These experiential details matter because they reveal not just what people think, but how they think and what truly motivates behaviour.
Synthetic respondents produce responses that sound credible because they have seen millions of similar conversations. But they do not have the underlying experiences that generate genuine insight. A synthetic persona representing a project manager might correctly list frustrations with collaboration tools, but cannot convey the visceral annoyance of a tool crashing during a client presentation because that is a lived moment, not a pattern in text.
This matters for research quality. When real participants struggle to articulate something or contradict themselves or share tangential stories, these messy human moments often contain the most valuable insights. Proper qualitative analysis depends on engaging with this complexity. Synthetic respondents give you clean, coherent responses devoid of the productive messiness that characterises real human communication.
Limitation 3: The sycophancy problem and calibrated emotions
One of the most discussed challenges with large language models is sycophancy - the tendency to agree with the user and provide overly positive responses. You can tune this behaviour in synthetic respondents, making them more critical or negative. But this creates a different problem: artificial grumpiness.
The fundamental issue is that human positivity and negativity are not parameters you can calibrate. They emerge from genuine experiences, preferences, and emotional responses. When someone says they love a product, that enthusiasm (or lack thereof) conveys real information about product-market fit, emotional engagement, and likelihood to recommend.
With synthetic respondents, you are choosing a positivity setting rather than measuring actual sentiment. If you tune them to be more critical to avoid sycophancy, you do not know whether their negativity reflects what real users would feel or just your calibration choices. If you leave them more positive, you cannot distinguish genuine enthusiasm from AI agreeableness.
This makes synthetic respondents particularly problematic for research on taste, emotional response, and buying intent. Would customers actually pay for this premium feature? Do users genuinely prefer design A over design B, or just find both acceptable? These questions require measuring real human preferences, not AI approximations of what preferences might look like.
The "non-bullshit answer" matters. Real research participants tell you when something is mediocre, unnecessary, or solving a problem they do not actually have. They express enthusiasm or indifference in ways that reveal true engagement levels. Synthetic respondents give you responses that fit expected patterns, not messy human truth.
Where synthetic respondents can be useful
Despite these limitations, synthetic respondents do have legitimate applications when used appropriately and transparently.
Testing interview guides and research instruments: Before fielding a study with real participants, you can use synthetic respondents to test whether your questions are clear, whether they elicit useful responses, and whether your interview flow makes sense. This is similar to piloting research, but faster and cheaper. Just remember that real pilots often reveal unexpected issues that synthetic testing misses.
Validating against known criteria for commoditised features: If you are building something that should match established standards, synthetic respondents can help verify completeness. Does your SaaS checkout flow include all the trust signals that successful competitors use? Does your product onboarding cover the typical steps users expect? For these checklist-style validations where the right answer is well-established, synthetic feedback can save time.
Exploring question phrasings and scenario variations: When designing complex research, you might want to test different ways of asking questions or presenting scenarios. Synthetic respondents let you rapidly iterate on research design without burning through participant pools. This is particularly useful for complex studies where question wording significantly impacts responses.
Where synthetic respondents fail: real research on new things
Synthetic respondents fundamentally cannot replace humans for research aimed at building genuine knowledge or understanding novel phenomena.
If you are exploring how people adapt to a genuinely new technology, understanding emerging customer needs in a changing market, investigating why a product is unexpectedly succeeding or failing, or seeking insights to drive innovation rather than imitation, you need real human participants. Full stop.
The most valuable research insights come from understanding what you do not already know. Synthetic respondents can only tell you variations on what is already in their training data. Real humans surprise you. They make connections you did not anticipate, express needs you had not considered, and reveal friction points that no existing research has documented.
Moreover, research is often about understanding not just what people think, but why they think it, how their thinking evolves, and what underlying experiences shape their perspectives. These deeper layers of understanding require engaging with real human complexity, not pattern-matched approximations.
For market researchers who need insights from real respondents, see how Skimle supports customer and market research workflows — including how AI-assisted analysis handles large batches of real interview transcripts and open-ended survey responses.
The real efficiency gains: interviewing and analysis
There is genuine irony in the synthetic respondents trend. The market research industry is pursuing a technology that saves time on recruitment and data collection (where the real work actually happens) while still leaving researchers with the hard parts: designing good research, asking the right questions, and making sense of responses.
The actual efficiency bottlenecks in qualitative research are not the interviews themselves. Most researchers would happily spend more time talking to real customers if they could. The pain points are:
- The time required to analyse interview transcripts systematically
- The weeks needed to code hundreds of open-ended survey responses
- The challenge of identifying patterns across dozens of conversations
- The difficulty of maintaining rigour while moving quickly
These are precisely the problems that AI can actually solve well in qualitative research when applied appropriately. Instead of replacing humans with synthetic approximations, use AI to make human insights more accessible and actionable.
Tools designed for proper qualitative analysis, like Skimle, can compress weeks of coding and analysis into days whilst maintaining methodological rigour. They can help you systematically categorise responses, identify patterns, and generate insights whilst preserving full transparency from every finding back to source data. This is where the real time savings and efficiency gains exist, not in replacing research participants with AI.
The bottom line: complement, not replacement
Synthetic respondents are a tool, not a revolution. Used transparently for appropriate applications like testing research instruments or validating against established standards, they can save time and money. Used as a replacement for real human research on genuine questions of customer understanding and innovation, they are a dangerous shortcut that produces plausible-sounding nonsense.
If you are considering synthetic respondents for your research, ask yourself: Am I trying to validate that I have met a known standard, or am I trying to learn something genuinely new? The first might be appropriate for synthetic feedback. The second requires real humans.
And if you are drowning in real human research data that you cannot analyse quickly enough, the solution is not synthetic respondents. The solution is better analysis tools that help you make sense of human insights faster.
Frequently asked questions
What are synthetic respondents in research?
Synthetic respondents are AI-generated personas that simulate how real people would respond to survey questions or interview prompts. They are created by training language models on demographic and behavioural data. They can be used for rapid hypothesis testing before fielding real research.
What are the main limitations of synthetic respondents?
Synthetic respondents reflect patterns in training data rather than genuine lived experience. They cannot capture genuinely novel opinions, recent events outside their training window, or the nuanced emotional responses that make qualitative research valuable. They are most reliable for predicting responses on well-established topics with abundant historical data.
When should you use synthetic respondents vs real interviews?
Use synthetic respondents for early-stage screening, questionnaire piloting, and testing obvious hypotheses quickly. Use real respondents whenever the research question requires genuine insight into behaviour, motivation, or experience — especially for new product categories, sensitive topics, or markets with limited historical data.
Can AI-generated respondents replace focus groups?
Not reliably. Focus groups surface unexpected social dynamics, minority opinions, and genuine emotional reactions that synthetic respondents cannot replicate. AI-generated respondents are better understood as a complement to real research rather than a replacement.
How do companies use synthetic respondents in practice?
Common uses include: generating a first-pass customer persona for a new market, pre-testing survey instruments before fielding, and simulating responses to messaging concepts. These applications work best when validated with at least a small sample of real respondents afterward.
Does new research prove synthetic respondents don't work?
Research published through 2026 shows they fail specifically at individual-level prediction and segment-targeting, the tasks most commercial research relies on. A cross-domain benchmark against the U.S. General Social Survey and World Values Survey found no large language model beat a simple demographic baseline, and that segment-targeting decisions built on synthetic data picked the wrong segment in around half of U.S. cases tested. That does not mean synthetic respondents are useless everywhere, but it does close off the higher-stakes uses vendors have been pitching them for.
Ready to analyse your qualitative research data with both speed and rigour? Try Skimle for free and experience systematic AI-assisted analysis with full transparency from every insight back to source data.
Want to learn more about proper qualitative analysis? Read our guides on how to analyse interview transcripts, thematic analysis methodology, and using AI responsibly in qualitative research.
Sources
- When Synthetic Users Fail: A Cross-Domain Benchmark of LLM-Simulated Human Survey Responses - Chen, Zhu and Zheng, arXiv (2026)
- Leaving Insight to Digital Twins? Promise, Progress and Limits of Synthetic Respondents - Kaiser et al., NIM Marketing Intelligence Review (2026)
- Synthetic Personalities: How Well Can LLMs Mimic Individual Respondents Using Socio-Economic Microdata? - Kinzinger and Hartmann, TUM School of Management, arXiv (2026)
- Synthetic Replacements for Human Survey Data? The Perils of Large Language Models - Political Analysis, Cambridge University Press (2025)
- Synthetic Data: Get on Board, But Do It Wisely - Research World
- Synthetic Respondents: The Promise, The Pitfalls, and The Reality Check - Merrill Research
- The Future of Synthetic Respondents in the Insights Industry - Quirks
- Rival Group's 2026 Market Research Trends Report
Further reading for market research teams: fake respondents in qualitative research: 8 warning signs.
About the author
Olli Salo is a former Partner at McKinsey & Company where he spent 18 years helping clients understand the markets and themselves, develop winning strategies and improve their operating models. He has done over 1000 client interviews and published over 10 articles on McKinsey.com and beyond. LinkedIn profile



