Benchmarking large language model agent societies against human behavioural distributions
SILICA evaluates LLM agent societies' behavior, finding alignment with human data only at initial stages.
Key Findings
Methodology
The paper introduces SILICA, a tool to assess whether LLM agent societies behave like humans. SILICA tests models in five environments, including repeated prisoner's dilemma and public goods games, using a perturbation library and fixed offer schedules to evaluate robustness and contamination.
Key Results
- Of 11 models, 8 matched human data in initial public goods contributions, but none matched end-state contributions.
- Only one model correctly set acceptance thresholds in fixed offer tests, others performed poorly.
- Reordering actions led to a 58-point drop in cooperation for one model.
Significance
This study provides a new tool for validating LLMs' effectiveness in simulated societies. By comparing with human behavior, it highlights models' limitations in complex social interactions, emphasizing caution in using these models for social science experiments.
Technical Contribution
SILICA offers a systematic approach to evaluate models' behavioral consistency and robustness. Through perturbation and contamination tests, it reveals behavioral changes under different conditions, providing new validation standards.
Novelty
SILICA is the first tool to evaluate multi-agent interactive simulations against human benchmarks, differing from previous single-response evaluations.
Limitations
- Models only align with human data initially, failing in later stages.
- Perturbation tests show models are sensitive to information presentation, affecting results.
Future Work
Future research should focus on improving models' consistency in complex social interactions and exploring broader perturbation conditions.
AI Executive Summary
Large language model (LLM) agent societies are being used for experimental simulations, but their effectiveness is questioned. The SILICA tool tests models' behavioral consistency and robustness across five environments, finding alignment with human data only at initial stages. Perturbation tests reveal models' sensitivity to information presentation, affecting behavioral consistency. The study emphasizes caution in using these models for social science experiments and provides new validation standards for future research.
The study shows significant limitations in models' performance in complex social interactions, especially when facing information changes. In fixed offer schedule tests, only one model correctly set acceptance thresholds, indicating deficiencies in handling complex decisions.
SILICA provides a new method for validating LLMs' effectiveness in simulated societies. Future research should focus on improving models' consistency in complex social interactions and exploring broader perturbation conditions to enhance applicability and reliability.
Deep Analysis
Background
Large language model (LLM) agent societies are used to simulate social behaviors, offering a lab without recruitment costs. However, their effectiveness is questioned, particularly in reflecting human behavior. Existing research focuses on single-response evaluations, lacking systematic validation of multi-agent interactive simulations.
Core Problem
The core issue is whether LLM agent societies can truly simulate human behavior, especially in complex social interactions. Consistency and robustness of models' behavior under different conditions are key challenges.
Innovation
SILICA introduces a perturbation library and fixed offer schedules to systematically evaluate multi-agent interactive simulations against human benchmarks. This innovation provides new standards for validating models' effectiveness.
Methodology
- �� SILICA tests five environments, including repeated prisoner's dilemma and public goods games.
- �� Uses a perturbation library to alter information presentation, evaluating behavioral changes.
- �� Fixed offer schedules test models' performance under different decision conditions.
Experiments
The experimental design includes five environment tests and application of the perturbation library. Fixed offer schedules evaluate models' acceptance threshold settings under different conditions. Experiments run on a single consumer-grade GPU, ensuring reproducibility.
Results
Results show models align with human data only initially, failing in later stages. Perturbation tests reveal models' sensitivity to information presentation, affecting behavioral consistency.
Applications
SILICA can be used to validate LLMs' effectiveness in social science experiments, helping researchers assess models' applicability in simulated societies.
Limitations & Outlook
Models show significant limitations in complex social interactions, especially when facing information changes. Future research should explore broader perturbation conditions to enhance models' applicability.
Plain Language Accessible to non-experts
Imagine a lab where intelligent robots are used to simulate human behavior. Researchers want to know if these robots can make decisions like humans. SILICA is like a testing instrument that helps researchers evaluate the robots' performance. By changing experimental conditions, such as how robots receive information, researchers can observe changes in robot behavior. It's like experimenting in a kitchen, adjusting different ingredients to see how the dish's taste changes. SILICA helps researchers find out why robots perform poorly and provides directions for future improvements.
ELI14 Explained like you're 14
Imagine playing a complex role-playing game with friends. The game has many characters, each with its own tasks and ways of making decisions. Researchers want to know if these characters can make decisions like real people. SILICA is like a super game tester that helps researchers evaluate these characters' performance. By changing game rules, such as how characters receive information, researchers can observe changes in character behavior. It's like adding new challenges to the game to see if the characters can adapt. SILICA helps researchers find out why characters perform poorly and provides directions for future improvements.
Glossary
Large Language Model
A deep learning-based model capable of generating and understanding natural language.
Used as agents in simulating social behaviors.
Perturbation Library
A set of tools for altering experimental conditions to test model robustness.
Used to evaluate models' performance under different conditions.
Fixed Offer Schedule
A testing method providing fixed offers to evaluate models' decision-making abilities.
Used to test models' performance under different decision conditions.
Behavioral Consistency
The ability of a model to maintain similar behavior under different conditions.
Evaluates models' effectiveness and robustness.
Data Contamination
Phenomenon where models exhibit abnormal behavior due to biases in training data.
Affects models' behavioral consistency and effectiveness.
Open Questions Unanswered questions from this research
- 1 How to improve models' consistency in complex social interactions? Current methods are sensitive to information changes, requiring more robust solutions.
- 2 How to reduce the impact of data contamination on results without affecting model performance?
Applications
Immediate Applications
Social Science Experiments
SILICA can be used to validate LLMs' effectiveness in social science experiments, helping researchers assess models' applicability in simulated societies.
Long-term Vision
Intelligent Agent Development
Improving models' consistency in complex social interactions can drive the development and application of intelligent agents, enhancing human-computer interaction experiences.
Abstract
Populations of large language model agents are increasingly used as experimental societies. Three doubts shadow every such result: whether the agents behave like the humans they stand in for, whether a finding survives changes to the apparatus that leave the rules untouched, and whether apparent social dynamics are interaction at all rather than the reproduction of experiments the models have read. This article introduces SILICA, an open instrument that tests all three. Five environments carry published human anchors, each paired with perturbations that re-render the same rules and with variants whose payoffs point away from the memorised result. Twelve open-weight models were run through it on a single consumer graphics card. Agreement with human data is confined to starting points: first-round public-goods contributions fall inside the equivalence margin for eight of eleven models, while no model matches end-state contributions or the human corridor of cooperation. Merely swapping the order in which two actions are listed costs one model 58 points of cooperation. Presenting responders with a fixed schedule of offers shows that only one model, the sole reasoning-trained one, places its acceptance threshold where the incentive requires; two move theirs part of the way, two move them the wrong way, and three never acquire one. Conventions form through a shared prior over the names rather than through negotiation, though negotiation reappears once that prior is disrupted. On the certification ladder defined here, current silicon societies support exploratory claims and no more.