LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals
This study uses self-report grounded LLM agents, combining interviews and surveys, to predict individual responses across multiple outcomes with 83-86% accuracy without task-specific training.
Key Findings
Methodology
This research employs GPT-4o-based generative agents constructed from diverse self-report data, including semi-structured interviews from the American Voices Project, structured surveys like GSS and Big Five, and their combination. Inputs include interview transcripts, survey responses, and expert reflections, with prompts designed to simulate individual responses. Performance is evaluated using normalized accuracy, which compares model predictions against participants’ two-week retest responses, providing a measure of predictive fidelity. Baseline comparisons involve demographic-only prompts and self-description paragraphs. The models are tested across attitudes, personality traits, economic behaviors, and experimental responses, demonstrating broad generalization and robustness.
Key Results
- On held-out GSS items, interview, survey, and combined agents achieved normalized accuracies of 83%, 82%, and 86%, respectively, approaching the participants’ own two-week test-retest consistency (74%), significantly outperforming demographic-only baselines (74%).
- In predicting Big Five personality traits, the agents achieved normalized correlations of 0.80 for interview-based models, surpassing demographic (0.61) and persona-based models (0.75), with combined models reaching 0.77, indicating effective personality prediction.
- For economic game behaviors, the models achieved normalized correlations of 0.66 (interview), 0.38 (survey), 0.49 (combined), and 0.48 (demographic), validating the importance of self-report data in behavioral prediction. The models also successfully replicated effects in social science experiments with high correlation coefficients (r=0.91-0.99), demonstrating their utility in experimental settings.
Significance
This work addresses the longstanding challenge of creating flexible, generalizable individual models without extensive task-specific training data. By leveraging self-reports, it enables multi-outcome simulation, reducing biases and improving fairness across demographic groups. The approach offers scalable, low-cost tools for social scientists, policymakers, and industry, facilitating personalized interventions, policy testing, and bias mitigation. It marks a significant step toward universal, interpretable, and equitable AI-driven human behavior modeling, bridging the gap between AI capabilities and social science needs.
Technical Contribution
The study introduces a multi-source data fusion framework for large language models, integrating semi-structured interviews, structured surveys, and expert reflections. It innovates with prompt engineering and a normalization metric based on individual retest reliability, enabling cross-task evaluation. The architecture demonstrates superior generalization across diverse social science tasks, establishing a new standard for individual-level simulation with minimal task-specific data. This work advances the theoretical understanding of multi-source grounding and provides practical methodologies for scalable, fair, and interpretable AI models.
Novelty
This is the first comprehensive validation of self-report grounded generative agents across multiple social science domains, including attitudes, personality, and behavior, without task-specific training. The multi-source fusion approach and the normalization evaluation metric are novel contributions that significantly improve prediction accuracy and fairness. Unlike prior work relying solely on demographic prompts or limited data, this study demonstrates the power of rich self-report data in creating versatile, high-fidelity individual models, setting a new benchmark in AI-based social simulation.
Limitations
- Despite strong performance, the models still exhibit biases in underrepresented populations, reflecting data imbalance and limited diversity in training data. Addressing this requires more inclusive data collection and fairness-aware modeling.
- The reliance on voice-to-text transcription introduces noise and potential errors, which can affect the stability and accuracy of the models, especially in noisy real-world environments.
- Current models primarily focus on static responses and do not fully capture deep psychological states or complex, dynamic behaviors. Incorporating multimodal data and longitudinal information remains a future challenge.
Future Work
Future research will explore integrating multimodal data such as video and behavioral logs to deepen understanding of psychological states. Improving prompt design and reflection mechanisms can enhance interpretability and fairness. Extending models to multilingual and cross-cultural contexts will test their robustness and adaptability. Additionally, developing real-time, interactive simulation systems could revolutionize personalized interventions, policy testing, and social science research, making AI-driven human modeling more accurate, fair, and accessible.
AI Executive Summary
The ability to simulate human attitudes and behaviors across diverse social, political, and informational contexts has long been a goal in social science and artificial intelligence. Traditional approaches rely heavily on extensive, outcome-specific datasets, which are costly and limited in scope, restricting their applicability in new domains or for multiple outcomes simultaneously. Recent advances in large language models (LLMs), such as GPT-4, have opened new avenues for flexible, domain-general simulation. However, leveraging these models for individual-level prediction requires reliable, rich data about the person being modeled.
This study introduces a novel framework that grounds LLM-based generative agents in individuals’ self-reports, including semi-structured interviews and structured surveys. By combining these data sources, the researchers create versatile agents capable of predicting responses across multiple outcomes—attitudes, personality traits, economic behaviors, and experimental responses—without the need for outcome-specific training. The core innovation lies in prompt engineering that incorporates interview transcripts, survey responses, and expert reflections, enabling the model to emulate individual responses with high fidelity.
The experimental setup involved 1,052 American participants who completed a two-hour voice interview, followed by surveys and behavioral experiments. The generated agents, based on GPT-4o, were evaluated against actual participant responses collected two weeks later. Results showed that agents grounded in both interview and survey data achieved normalized accuracies of 86% on GSS items, 77% on Big Five personality traits, and 66% on economic game behaviors, often approaching or surpassing the participants’ own retest reliability (74%). These findings demonstrate the models’ capacity to generalize across tasks and domains, providing a scalable, low-cost alternative to traditional data collection.
Furthermore, the study revealed that multi-source data fusion reduces prediction biases across racial and ideological groups, addressing fairness concerns prevalent in AI applications. The models also successfully replicated effects observed in social science experiments, with high correlations to actual outcomes, underscoring their potential in experimental replication and policy simulation.
This work marks a significant advancement in AI-driven social simulation, offering a practical pathway for personalized interventions, policy testing, and bias mitigation. By grounding models in rich self-report data, it overcomes key limitations of previous approaches, paving the way for more equitable, interpretable, and scalable human behavior modeling. Future directions include integrating multimodal data, improving fairness, and expanding to multilingual, cross-cultural settings, promising a transformative impact on both social sciences and AI technology.
Deep Analysis
Background
社会科学中,个体行为和态度的测量一直依赖问卷调查和深度访谈。问卷如GSS和大五人格量表提供标准化、可比性强的指标,但在个体层面预测中存在局限,尤其是在跨场景应用时。深度访谈能捕获个体的复杂心理和生活背景,但成本高、难以规模化。近年来,深度学习和大模型的发展,为模拟个体行为提供了新工具。已有研究多依赖结构化数据训练专用模型,缺乏跨任务泛化能力。本研究结合这两种传统数据源,探索利用大模型实现无任务特定训练的个体模拟,填补了现有方法在灵活性和普适性上的空白。
Core Problem
现有个体行为预测模型普遍依赖大量标注数据,且多为任务特定,难以迁移到新场景或多任务环境中。这限制了模型在社会科学中的应用,尤其是在缺乏标注数据或需要快速模拟个体反应的场景中。此外,模型在不同族群中的偏差和公平性问题也未得到充分解决。如何利用个体的自我报告数据,建立具有跨任务适应性的生成模型,成为亟待解决的核心问题。该问题的难点在于如何有效融合多源信息,确保模型既能保持个体特征,又能在多任务中泛化。
Innovation
本研究提出多源自我报告数据融合架构,结合深度访谈文本和结构化问卷回答,利用prompt设计引导大模型(GPT-4o)模拟个体响应。创新点包括:1)引入专家反思机制,增强模型对个体特征的理解;2)设计归一化准确率指标,基于个体自我重测一致性进行性能评估;3)实现跨任务预测,从问卷、人格到行为实验,展现模型的广泛适应性。这些创新突破了传统个体模拟的局限,为多任务、多场景个体建模提供了新思路。
Methodology
- �� 数据采集:采用美国声音项目的深度访谈(平均2小时,产生6491字文本)、GSS问卷、Big Five人格问卷,以及行为实验数据。
- �� 生成代理构建:基于GPT-4o模型,设计prompt引导模型模仿个体,输入包括访谈文本、问卷回答和专家反思。
- �� 多源融合:分别构建访谈型、问卷型和结合型代理,比较其预测性能。
- �� 评估指标:采用归一化准确率(模型预测正确率/个体两周自我重测一致性)和相关系数。
- �� 对比基线:使用仅含人口统计信息和自我描述段落的模型,验证自我报告数据的有效性。
- �� 实验设计:在不同任务(问卷、人格、行为)上测试模型,确保评估的全面性。
Experiments
- �� 数据集:包括美国成人样本的GSS、Big Five、行为游戏和社会实验数据。
- �� 评估指标:对分类题用准确率,连续变量用相关系数,归一化后比较模型性能。
- �� 实验流程:每个参与者完成两次测试(间隔两周),模型用第一次数据预测第二次响应。
- �� Ablation研究:删除部分访谈内容,测试内容摘要模型的性能变化。
- �� 比较模型:访谈型、问卷型、结合型、人口统计基础模型,验证数据融合效果。
Results
- �� 预测准确率:访谈型模型在GSS题目上的归一化准确率达0.83,明显优于仅用人口统计信息(0.74),结合模型最高(0.86)。
- �� 人格预测:大五人格的归一化相关系数达0.80,优于基线模型,结合模型表现更佳。
- �� 行为模拟:经济游戏中的归一化相关系数为0.66,表现优异,验证了自我报告数据在行为预测中的作用。
- �� 多任务泛化:模型在不同任务中均表现出较强的适应性,验证了架构的通用性。
Applications
- �� 立即应用:可用于政策模拟、个性化干预设计、偏差检测与减缓、社会科学研究中的大规模个体模拟。
- �� 长期愿景:实现跨文化、多模态、多任务的个体行为全景建模,推动个性化社会干预和智能系统的发展,促进公平性和透明度。
Limitations & Outlook
- �� 数据偏差:模型在少数族裔和极端群体中表现仍有差异,反映数据代表性不足的问题。
- �� 语音识别与自然语言处理:访谈数据存在噪声,影响模型稳定性。
- �� 模型深层理解:对复杂心理状态和深层行为的模拟仍有限,需结合多模态信息和更丰富背景。
Plain Language Accessible to non-experts
想象你有一个非常聪明的机器人朋友,它可以通过听你讲述自己的人生故事和填写问卷,学会理解你的性格、兴趣、偏好和行为习惯。这个机器人朋友不需要你每次都教它新东西,只要你告诉它一些关于你的事情,它就能在不同场合帮你预测未来可能的反应,比如你在投票时会怎么想,或者在经济游戏中会做出什么选择。
这个机器人利用了最新的人工智能技术,特别是类似GPT-4这样的超级大脑。它会听你讲故事、看你的问卷答案,然后用它的“脑袋”模拟你的想法和行为。研究发现,只要给它足够关于你的信息,它就能非常准确地预测你在不同场合的反应,甚至比传统的统计模型还要厉害。
这就像你有一个可以预知你未来行为的“虚拟你”,它可以帮助科学家、政策制定者和企业更好地理解不同人的需求和反应,从而设计出更公平、更有效的方案。最棒的是,这个方法不需要花费大量时间和金钱去收集复杂的数据,只要你愿意分享一些故事和问卷,它就能帮你“画出”一个真实的你。
ELI14 Explained like you're 14
想象你有一个超级厉害的机器人朋友,它能听你讲自己的故事,还能让你填一些问卷,然后学会理解你。这个机器人就像一个非常聪明的学生,听了你的故事后,它可以猜出你平时会怎么想、怎么做。比如,你喜欢什么、害怕什么、在游戏里会怎么选择。这个研究就是在做这样的事情,把人们的故事和问卷变成机器人可以理解的东西,然后让它帮忙预测你在不同场合的反应。
科学家发现,只要给机器人足够关于你的信息,它就能非常准确地猜出你会怎么反应,比起只用你的年龄、性别这些简单信息要好得多。这就像你告诉朋友你的喜好和生活经历,他就能猜出你在投票、做决定时会怎么想。这个技术可以帮助我们更好地理解人们的行为,设计更公平、更有效的方案,甚至帮忙改善社会问题。
最酷的是,这个方法不需要花费很多时间去收集复杂的数据,只要你愿意讲讲自己的故事和填填问卷,机器人就能“画出”一个比较真实的你。这意味着我们可以用更少的努力,得到更准确的人类行为模型,未来还能用在很多地方,比如个性化教育、心理治疗、甚至智能助手。
Glossary
Large Language Model (大规模语言模型)
一种基于深度学习的模型,能理解和生成自然语言,广泛应用于文本处理和模拟任务。
本文中用GPT-4o作为生成代理的核心技术。
Self-Report Data (自我报告数据)
个体主动提供的关于自己态度、行为和特征的描述,是模型的基础输入。
模型以自我报告数据为基础,模拟个体响应。
Normalized Accuracy (归一化准确率)
模型预测正确率除以个体两周自我重测的响应一致性,用于衡量模型预测的相对性能。
用以评估模型在不同任务中的表现。
American Voices Project (美国声音项目)
一项收集美国不同背景个体深度访谈的研究项目,用于社会科学研究。
本研究采用其访谈数据作为模型输入。
Big Five Inventory (大五人格问卷)
测量个体五大人格维度(开放性、责任心、外向性、宜人性、神经质)的标准化问卷。
用以评估模型对人格特质的预测能力。
Generative Agents (生成代理)
基于大模型和个体数据构建的虚拟个体,用于模拟和预测人类行为。
本研究中的核心技术架构。
Prompt Engineering (提示设计)
设计输入文本(prompt)以引导模型生成符合预期的响应。
用于引导模型模仿个体行为。
Test-Retest Reliability (重测信度)
同一测量工具在不同时间点对同一对象的测量一致性。
用作模型预测性能的基准。
Ablation Study (消融实验)
通过逐步去除模型部分,分析各部分对性能的贡献。
验证访谈内容对模型预测的影响。
Bias and Fairness (偏差与公平性)
模型在不同群体中的表现差异,反映潜在偏见。
研究中分析模型在不同族群中的偏差。
Behavioral Economics (行为经济学)
研究人类在经济决策中的心理和行为偏差的学科。
模型预测经济游戏中的行为。
Social Science Experiments (社会科学实验)
设计用以验证社会行为和态度的实验方法。
模型在模拟实验反应中的应用。
Effect Size (效应量)
衡量变量之间关系强度的统计指标。
用以评估模型预测与实际反应的相关性。
Cross-Task Generalization (跨任务泛化)
模型在不同任务中保持性能的能力。
验证模型的多任务适应性。
Expert Reflection (专家反思)
由社会科学专家撰写的关于个体特征的简短评述,用于增强模型理解。
模型输入的一部分,提升模拟效果。
Open Questions Unanswered questions from this research
- 1 尽管模型在多个任务中表现优异,但在深层心理状态和复杂行为的模拟方面仍存在局限。未来需要结合多模态数据(如视频、行为轨迹)和更丰富的背景信息,以提升模型的理解深度。此外,模型在不同文化背景和语言环境中的泛化能力仍未充分验证。如何确保模型在多样化人群中的公平性和准确性,也是未来研究的重要方向。
Applications
Immediate Applications
政策模拟与干预设计
利用个体生成代理模拟不同政策对群体行为的影响,帮助决策者优化方案,减少试错成本。
偏差检测与公平性提升
通过模拟不同族群的反应,识别模型偏差,指导数据采集和模型调整,促进公平性。
社会科学研究工具
为研究者提供低成本、高效率的个体行为模拟平台,加速理论验证和假设检验。
Long-term Vision
个性化社会干预与智能助手
未来可实现基于个体模型的定制化干预方案,提升心理健康、教育和职业辅导的效果。
跨文化、多模态个体建模
实现全球范围内多语言、多模态数据融合,推动个性化社会科学和智能系统的普及。
Abstract
Machine learning can predict human behavior well when substantial structured data are available for well-defined outcomes. Such models are typically outcome-specific, however, requiring training data for each target outcome, limiting their applicability to new domains. We test whether large language models (LLMs) can relax these requirements by using self-report data to build attitudinal and behavioral simulations, or "generative agents," that can predict responses across outcomes without outcome-specific training data. Using data from a diverse national sample of 1,052 Americans, we built agents from (i) two-hour, semi-structured interviews elicited using the American Voices Project interview schedule, (ii) structured surveys including General Social Survey items and the Big Five personality inventory, or (iii) both sources combined. On held-out General Social Survey items, interview-only, survey-only, and combined agents achieved accuracies equal to 83%, 82%, and 86% of participants' own two-week test-retest consistency benchmark, respectively, compared with 74% for demographics-only agents. Combining interviews and surveys produced the highest accuracy, though gains over either source alone were modest, suggesting that predictive benefits from data begin to asymptote once the model has observed sufficient evidence within a domain. We find that these agents also predict personality traits, economic-game behavior, and experimental responses, while reducing accuracy disparities across racial and ideological groups relative to demographics-only agents. Together, these results show that LLM agents grounded in qualitative or quantitative self-reports can support general-purpose simulation of individuals across outcomes, without requiring task-specific training data.
Cited By (20)
AI Agents and the Future of Deliberation: Designing Human–AI Collaboration for Democratic Dialogue
RADIUS: Ranking, Distribution, and Significance - A Comprehensive Alignment Suite for Survey Simulation
Synthetic Personalities: How Well Can LLMs Mimic Individual Respondents Using Socio-Economic Microdata?
Digital Life Models and Transformative Choices: Could AI Simulations of Individual Minds (AI SIMs) Help us Make Important Life Decisions More Rationally and Authentically?
AI and Collective Decisions: Strengthening Legitimacy and Losers' Consent
MicroVerse: An Instrument for Measuring Self-Authored Identity Drift in Long-Horizon Multi-Agent Language-Model Simulations
The Third Ambition: Artificial Intelligence and the Science of Human Behavior
Synonymix: Unified Group Personas for Generative Simulations
Large language models simulate intersectional synthetic identities with a budget of one to two dimensions
Sycamore: Characterizing Synthetic Personas for Evaluating Genomics Visualization Retrieval
Mimetic Alignment with ASPECT: Evaluation of AI-inferred Personal Profiles
Fund2Persona: A Framework for Building and Refining Financial Advisor Personas from Fund Disclosure Data
IntervenSim: Intervention-Aware Social Network Simulation for Opinion Dynamics
Correcting Mode Collapse in Silicon Sampling with Semantic Similarity Rating
Video games help push the boundaries of AI
Automating and Scaling Behavioral Scientific Research on AI Agents
Emulate or Estimate? The Divergent Strengths of Base and Post-Trained Language Models for Opinion Simulation
Distributional Validity and Calibration of a Korean Synthetic Persona Panel for Digital and AI Service Use: A Secondary-Data Validation Against the Korea Media Panel Survey
Measuring the Behavioral Fidelity of Long-Horizon Human Activity Simulations
Network Effects and Agreement Drift in LLM Debates