Interview-Informed Generative Agents for Product Discovery: A Validation Study
A memory–retrieval–reflection agent study with 51 users and 4 concepts found distributional calibration but weak identity fidelity.
Key Findings
Methodology
The authors conducted 90-minute workflow interviews with 51 knowledge workers and used each transcript to instantiate a personalized LLM agent. The agents followed a memory–retrieval–reflection architecture: interview experiences were stored, relevant memories retrieved, and reflections used to generate concept evaluations. Agents and participants then assessed four AI document-workflow concepts using TAM, NPS, and open-ended feedback.
Key Results
- Agents approximated human response distributions and outperformed the reported baseline approaches at the population level. However, the supplied text gives no correlation, error, significance, or baseline values, so the evidence supports distributional calibration but not a specific percentage improvement.
- Individual fidelity was limited: an agent did not reliably reproduce the ratings, preferences, or rationales of the participant whose interview grounded it. The central pattern is therefore “identity-imprecise” simulation despite population-level similarity.
- The study contained 51 × 4 × 15 = 3,060 human responses and an equal number of simulated responses. Participants contributed about 43 minutes of raw interview audio on average, with responses averaging 106 words; a three-day retest measured human self-consistency.
Significance
The paper extends LLM simulation from established social-science instruments to novel products that users have never experienced. It separates two forms of validity often conflated: matching a population distribution and reproducing a particular person. For industry, this suggests a low-cost tool for early screening and directional comparison, but not a substitute for interviews about trust, workflow nuance, adoption barriers, or individual risk perceptions.
Technical Contribution
The work grounds a memory–retrieval–reflection agent in rich workflow interviews and validates it against the same participants’ actual concept-test responses. It evaluates both six-item TAM ratings and NPS, plus open-ended rationales, linking scalar prediction to contextual design feedback. The contribution is not a new foundation model or formal guarantee; it is an empirical characterization of a product-discovery boundary: interview grounding can improve population-level calibration without ensuring person-level behavioral reproduction.
Novelty
Relative to Park et al.’s evaluation on standardized instruments such as General Social Survey items, this is presented as the first systematic validation of interview-informed agents for early-stage product concept testing. The key novelty is testing preferences constructed on the fly for unfamiliar AI workflows while assessing both quantitative acceptance measures and qualitative design feedback.
Limitations
- The sample contains only 51 U.S.-based knowledge workers who use documents daily, limiting generalization to consumers, other cultures, less document-heavy occupations, and non-Western work practices.
- The supplied paper text does not report exact TAM/NPS errors, correlations, confidence intervals, or baseline scores. Consequently, the strength of the quantitative advantage and statistical reliability cannot be independently assessed from this excerpt.
- All four concepts concern AI document workflows; transfer to healthcare, education, consumer products, or high-stakes automation remains untested.
Future Work
Future studies should broaden cultural, occupational, and product coverage; preregister identity- and distribution-level fidelity thresholds; and compare foundation models, retrieval policies, and reflection prompts. Longitudinal deployment should test whether agents predict actual trial, retention, payment, and task behavior rather than stated concept attitudes. Responsible systems also need consent, privacy, withdrawal, provenance, and human-review controls.
AI Executive Summary
Product teams routinely ask people to judge products that do not yet exist. Conventional interviews provide rich insight but are slow and expensive, while earlier LLM simulations have focused mainly on established surveys and personality instruments. Wang and Siu ask whether agents grounded in real workflow interviews can simulate responses to unfamiliar product concepts.
The study recruited 51 U.S. knowledge workers for 90-minute interviews, then tested four AI document concepts: a Multi-document Q&A Assistant, Smart Highlights Assistant, Audio Assistant, and Workflow Actions Assistant. Each concept was assessed with six TAM items, NPS, and open-ended feedback. The dataset comprised 3,060 human responses, matched by an equal number of agent responses. Agents used a memory–retrieval–reflection architecture to draw on personal workflows, pain points, and technology-adoption experiences.
The result is a precise boundary rather than a simple success story. Agents were distribution-calibrated: their aggregate responses approximated human population patterns and outperformed the reported baselines. Yet they were identity-imprecise: they could not reliably reproduce the particular participant’s scores or explanations. The practical implication is that simulation may accelerate early concept screening and trade-off exploration, but authentic users remain essential for trust, workflow details, and adoption barriers. The paper’s main contribution is a validation framework and a caution against confusing population resemblance with personal replication.
Deep Analysis
Background
Prior work has shown that LLMs can simulate some survey and personality responses. Park et al., for example, reported 85% normalized accuracy on General Social Survey items relative to human two-week test–retest consistency. Product discovery is harder: users evaluate unfamiliar artifacts, construct preferences in real time, and reason about expected utility, effort, risk, and workflow fit rather than reporting established attitudes.
Core Problem
The study asks two questions. RQ1 tests whether interview-informed agents reproduce scalar TAM and NPS responses at individual and population levels. RQ2 examines how agent-generated open-ended feedback resembles or diverges from participants’ own comments. The bottleneck is converting contextual interview evidence into stable, testable predictions about hypothetical products.
Innovation
- �� It transfers interview-informed agents from standardized social-science instruments to early product discovery. • It combines scalar acceptance measures with qualitative design rationales. • It uses each participant’s own concept-test responses as ground truth. • It frames fidelity in two dimensions—distributional calibration and identity precision—rather than treating “accuracy” as a single quantity.
Methodology
- �� Screening: recruit workers using PDFs at least weekly in Finance, Legal, Operations, Management, or Research. • Interview: collect a 90-minute asynchronous audio account of workflows, pain points, AI experience, and adoption patterns. • Agent construction: organize transcript content as retrievable memory and generate responses through retrieval and reflection. • Evaluation: present four concepts and 15 questions per concept, including six 7-point Likert TAM items and NPS. • Validation: compare human responses, agent responses, baselines, and three-day human retest consistency; inspect open-ended workflow fit, concerns, and suggestions.
Experiments
The sample included 51 participants aged 21–76 years (M = 43.2, SD = 11.9), with 26 men and 25 women. Participants completed an interview, initial concept testing within 24 hours, and identical retesting three days later. Concepts spanned passive highlighting, audio interaction, multi-document synthesis, and active workflow execution. The study collected 3,060 human responses plus matched simulations and roughly 36 hours of speech. The excerpt does not specify the foundation model, prompts, hyperparameters, or complete baseline definitions.
Results
The headline result is strong population resemblance but weak person-level matching. Agents approximated aggregate human distributions and exceeded the reported baselines, yet did not reliably recover the target individual’s ratings or rationales. Responses averaged 106 words, and interviews averaged about 43 minutes of raw audio per participant. Because exact TAM/NPS errors, correlations, and significance tests are absent from the supplied text, no additional numerical improvement should be inferred.
Applications
Product teams can use agents to screen many concepts cheaply, rank directions, explore feature trade-offs, and identify recurring concerns before investing in prototypes. The prerequisite is a consented, sufficiently rich interview corpus and a human validation stage. In domains involving privacy, regulation, safety, trust, or consequential automation, simulations should be triage evidence only, never the sole basis for launch or design decisions.
Limitations & Outlook
External validity is constrained by a small, U.S.-only, document-intensive sample and four related concepts. Agents may compress complex interviews into stereotypes or produce unusually polished, unrealistic rationales. The study also leaves implementation details and exact metrics underspecified in the supplied text. Future work should test cross-cultural and cross-domain transfer, publish full evaluation protocols, and connect simulation to observed adoption, retention, productivity, and economic outcomes.
Plain Language Accessible to non-experts
Imagine a restaurant preparing four new dishes. Before cooking them for everyone, the chef interviews 51 customers about what they usually eat, what they avoid, and how adventurous they are. The chef then creates a digital stand-in for each customer. Each stand-in reads that customer’s food story and predicts whether the four unfamiliar dishes would be useful, easy, recommendable, or worth improving.
The stand-ins work reasonably well if the question is, “Which dish will this whole group probably like?” Their combined answers resemble the group’s overall pattern. But they are much less reliable when asked, “What exactly will Customer 17 think, and why?” A stand-in may remember that the customer dislikes complicated meals, yet still miss a personal worry or a strange situation that changes the decision.
That is the paper’s central lesson. Digital stand-ins can help a restaurant remove obviously weak dishes and compare many ideas quickly. They cannot replace the customer when the chef needs to understand a specific allergy, cultural habit, or hidden reason for rejection. The AI is a useful map of the crowd, not a perfect copy of one person.
ELI14 Explained like you're 14
Picture a team designing a new study app before the app even exists. They interview 51 students about homework habits, annoying parts of school, and whether they like automatic help. Then a computer reads each interview and creates a “digital twin” for every student. The twins judge four imaginary features: finding answers across many documents, highlighting important parts, turning reading into audio, and automatically completing a chain of tasks.
Researchers compare the real students’ answers with the twins’ answers. Here is the surprising part: when they look at the whole class, the twins are fairly good at spotting general trends. They can help show which kind of feature sounds popular. But if they ask, “What will Student 3 personally score this, and what exact worry will they mention?” the twin becomes shaky.
It is like a video-game character with your profile, favorite gear, and old scores—but not your real feelings when a level suddenly gets difficult. The character can guess your usual style, not perfectly recreate every decision. So AI twins are useful for quickly sorting ideas, but humans still need to explain trust, embarrassment, confusion, and “no way, I would never use that!”
The big message is not that AI is useless. It is that different jobs need different evidence: use simulations as a fast radar for group trends, then ask real people before making important choices!
Glossary
Generative agent
An LLM-driven system that generates simulated behavior or responses from contextual information. It models a possible participant response rather than being the participant.
The paper creates one personalized agent per interview participant.
Memory–retrieval–reflection architecture
A design that stores experiences, retrieves relevant memories, and uses reflection to organize a response. It connects generation to a simulated person’s background.
This architecture grounds concept evaluations in the 90-minute workflow interviews.
Technology Acceptance Model (TAM)
A framework measuring perceived usefulness, perceived ease of use, and behavioral intention. Here, each construct uses two items on a 7-point Likert scale.
TAM supplies the main scalar measures for human–agent comparison.
Net Promoter Score (NPS)
A recommendation-intent measure used as an indicator of advocacy or satisfaction. It complements TAM by capturing a broader willingness to recommend.
A standard NPS question was included for every concept.
Distributional calibration
Agreement between simulated and human response distributions at the group level. It does not imply that each individual is reproduced accurately.
It is the paper’s central description of aggregate agent performance.
Identity precision
The ability to reproduce the distinctive ratings, preferences, and rationales of a particular person. Rich grounding does not automatically provide identity-level replication.
The agents were described as identity-imprecise.
Open Questions Unanswered questions from this research
- 1 The supplied text omits exact errors, correlations, significance tests, and baseline values. It therefore remains unclear how strong calibration is for each TAM construct, NPS, or qualitative category, and which agent component contributes most.
- 2 The study does not establish whether simulated concept reactions predict real trial, continued use, payment, productivity, or retention. Longitudinal behavioral validation is needed.
- 3 The sample is restricted to U.S. knowledge workers. Cross-cultural validity, demographic bias, privacy, consent, and performance in high-stakes domains remain open questions.
Applications
Immediate Applications
Early concept screening
Teams can use interview-grounded agents to compare many AI concepts, inspect aggregate TAM/NPS tendencies, and surface recurring concerns before building prototypes. Human interviews and prototype tests should confirm any decision that affects launch, safety, trust, or compliance.
Directional design iteration
Designers can ask agents to compare workflow fit, implementation concerns, use cases, and improvements across variants. This supports brainstorming and prioritization, but agent rationales should be treated as directional hypotheses rather than authentic participant testimony.
Long-term Vision
Human–agent research infrastructure
A future platform could connect consented interview memories, large-scale agent exploration, and repeated human calibration. Agents would widen the search space while participants validate critical claims, with provenance, privacy, withdrawal, and audit controls built in.
Abstract
Large language models (LLMs) have shown strong performance on standardized social science instruments, but their value for product discovery remains unclear. We investigate whether interview-informed generative agents can simulate user responses in concept testing scenarios. Using in-depth workflow interviews with knowledge workers, we created personalized agents and compared their evaluations of novel AI concepts against the same participants' responses. Our results show that agents are distribution-calibrated but identity-imprecise: they fail to replicate the specific individual they are grounded in, yet approximate population-level response distributions. These findings highlight both the potential and the limits of LLM simulation in design research. While unsuitable as a substitute for individual-level insights, simulation may provide value for early-stage concept screening and iteration, where distributional accuracy suffices. We discuss implications for integrating simulation responsibly into product development workflows.