Benchmarking large language model agent societies against human behavioural distributions

TL;DR

SILICA evaluates LLM agent societies' behavior, finding alignment with human data only at initial stages.

physics.soc-ph 🔴 Advanced 2026-08-28 6 views
Raad Bin Tareaf
large language models social simulation behavioral validation data contamination experimental sociology

Key Findings

Methodology

The paper introduces SILICA, a tool to assess whether LLM agent societies behave like humans. SILICA tests models in five environments, including repeated prisoner's dilemma and public goods games, using a perturbation library and fixed offer schedules to evaluate robustness and contamination.

Key Results

  • Of 11 models, 8 matched human data in initial public goods contributions, but none matched end-state contributions.
  • Only one model correctly set acceptance thresholds in fixed offer tests, others performed poorly.
  • Reordering actions led to a 58-point drop in cooperation for one model.

Significance

This study provides a new tool for validating LLMs' effectiveness in simulated societies. By comparing with human behavior, it highlights models' limitations in complex social interactions, emphasizing caution in using these models for social science experiments.

Technical Contribution

SILICA offers a systematic approach to evaluate models' behavioral consistency and robustness. Through perturbation and contamination tests, it reveals behavioral changes under different conditions, providing new validation standards.

Novelty

SILICA is the first tool to evaluate multi-agent interactive simulations against human benchmarks, differing from previous single-response evaluations.

Limitations

  • Models only align with human data initially, failing in later stages.
  • Perturbation tests show models are sensitive to information presentation, affecting results.

Future Work

Future research should focus on improving models' consistency in complex social interactions and exploring broader perturbation conditions.

AI Executive Summary

Large language model (LLM) agent societies are being used for experimental simulations, but their effectiveness is questioned. The SILICA tool tests models' behavioral consistency and robustness across five environments, finding alignment with human data only at initial stages. Perturbation tests reveal models' sensitivity to information presentation, affecting behavioral consistency. The study emphasizes caution in using these models for social science experiments and provides new validation standards for future research.

The study shows significant limitations in models' performance in complex social interactions, especially when facing information changes. In fixed offer schedule tests, only one model correctly set acceptance thresholds, indicating deficiencies in handling complex decisions.

SILICA provides a new method for validating LLMs' effectiveness in simulated societies. Future research should focus on improving models' consistency in complex social interactions and exploring broader perturbation conditions to enhance applicability and reliability.

Deep Analysis

Background

Large language model (LLM) agent societies are used to simulate social behaviors, offering a lab without recruitment costs. However, their effectiveness is questioned, particularly in reflecting human behavior. Existing research focuses on single-response evaluations, lacking systematic validation of multi-agent interactive simulations.

Core Problem

The core issue is whether LLM agent societies can truly simulate human behavior, especially in complex social interactions. Consistency and robustness of models' behavior under different conditions are key challenges.

Innovation

SILICA introduces a perturbation library and fixed offer schedules to systematically evaluate multi-agent interactive simulations against human benchmarks. This innovation provides new standards for validating models' effectiveness.

Methodology

  • �� SILICA tests five environments, including repeated prisoner's dilemma and public goods games.
  • �� Uses a perturbation library to alter information presentation, evaluating behavioral changes.
  • �� Fixed offer schedules test models' performance under different decision conditions.

Experiments

The experimental design includes five environment tests and application of the perturbation library. Fixed offer schedules evaluate models' acceptance threshold settings under different conditions. Experiments run on a single consumer-grade GPU, ensuring reproducibility.

Results

Results show models align with human data only initially, failing in later stages. Perturbation tests reveal models' sensitivity to information presentation, affecting behavioral consistency.

Applications

SILICA can be used to validate LLMs' effectiveness in social science experiments, helping researchers assess models' applicability in simulated societies.

Limitations & Outlook

Models show significant limitations in complex social interactions, especially when facing information changes. Future research should explore broader perturbation conditions to enhance models' applicability.

Plain Language Accessible to non-experts

Imagine a lab where intelligent robots are used to simulate human behavior. Researchers want to know if these robots can make decisions like humans. SILICA is like a testing instrument that helps researchers evaluate the robots' performance. By changing experimental conditions, such as how robots receive information, researchers can observe changes in robot behavior. It's like experimenting in a kitchen, adjusting different ingredients to see how the dish's taste changes. SILICA helps researchers find out why robots perform poorly and provides directions for future improvements.

ELI14 Explained like you're 14

Imagine playing a complex role-playing game with friends. The game has many characters, each with its own tasks and ways of making decisions. Researchers want to know if these characters can make decisions like real people. SILICA is like a super game tester that helps researchers evaluate these characters' performance. By changing game rules, such as how characters receive information, researchers can observe changes in character behavior. It's like adding new challenges to the game to see if the characters can adapt. SILICA helps researchers find out why characters perform poorly and provides directions for future improvements.

Glossary

Large Language Model

A deep learning-based model capable of generating and understanding natural language.

Used as agents in simulating social behaviors.

Perturbation Library

A set of tools for altering experimental conditions to test model robustness.

Used to evaluate models' performance under different conditions.

Fixed Offer Schedule

A testing method providing fixed offers to evaluate models' decision-making abilities.

Used to test models' performance under different decision conditions.

Behavioral Consistency

The ability of a model to maintain similar behavior under different conditions.

Evaluates models' effectiveness and robustness.

Data Contamination

Phenomenon where models exhibit abnormal behavior due to biases in training data.

Affects models' behavioral consistency and effectiveness.

Open Questions Unanswered questions from this research

  • 1 How to improve models' consistency in complex social interactions? Current methods are sensitive to information changes, requiring more robust solutions.
  • 2 How to reduce the impact of data contamination on results without affecting model performance?

Applications

Immediate Applications

Social Science Experiments

SILICA can be used to validate LLMs' effectiveness in social science experiments, helping researchers assess models' applicability in simulated societies.

Long-term Vision

Intelligent Agent Development

Improving models' consistency in complex social interactions can drive the development and application of intelligent agents, enhancing human-computer interaction experiences.

Abstract

Populations of large language model agents are increasingly used as experimental societies. Three doubts shadow every such result: whether the agents behave like the humans they stand in for, whether a finding survives changes to the apparatus that leave the rules untouched, and whether apparent social dynamics are interaction at all rather than the reproduction of experiments the models have read. This article introduces SILICA, an open instrument that tests all three. Five environments carry published human anchors, each paired with perturbations that re-render the same rules and with variants whose payoffs point away from the memorised result. Twelve open-weight models were run through it on a single consumer graphics card. Agreement with human data is confined to starting points: first-round public-goods contributions fall inside the equivalence margin for eight of eleven models, while no model matches end-state contributions or the human corridor of cooperation. Merely swapping the order in which two actions are listed costs one model 58 points of cooperation. Presenting responders with a fixed schedule of offers shows that only one model, the sole reasoning-trained one, places its acceptance threshold where the incentive requires; two move theirs part of the way, two move them the wrong way, and three never acquire one. Conventions form through a shared prior over the names rather than through negotiation, though negotiation reappears once that prior is disrupted. On the certification ladder defined here, current silicon societies support exploratory claims and no more.

physics.soc-ph cs.CL