Message capacity and claim wording set the transition points of collective truth-finding in language-model networks

TL;DR

Study explores how message capacity and claim wording set transition points in language-model networks.

cs.MA 🔴 Advanced 2026-09-15 8 views
Makoto Fukushima
language model message capacity collective decision truth-finding network model

Key Findings

Methodology

The study models message capacity to determine how many messages an agent reads and generates communication networks. It analyzes an 8B parameter model's judgment using a stochastic binary neuron update rule.

Key Results

  • In 31,824 queries, model judgment reduces to a logistic function of its inbox's weighted sum.
  • In 1,414 episodes, the correct side won only 28-45% of the time.
  • Assertion bias not detected in 70B model.

Significance

The study reveals how two single-agent measurements—claim wording threshold and message capacity—largely determine collective fate in decision-making.

Technical Contribution

Predicts collective transition points from single-agent measurements, offering a new theoretical framework for analyzing collective behavior in language-model networks.

Novelty

First to combine message capacity and claim wording to predict collective truth-finding in language-model networks.

Limitations

  • In some experiments, predictions failed with correct side winning less than 50%.
  • Claim wording threshold failed to transfer across different claim sets.

Future Work

Future research could explore larger-scale model behavior and the impact of different claim wordings on collective decisions.

AI Executive Summary

The paper investigates how message capacity and claim wording influence transition points in language-model networks. By modeling the number of messages an agent reads, generating communication networks, and analyzing an 8B parameter model's judgment process, the study shows that wrong consensus becomes unreachable when agents read fewer than 6.4 messages on average. However, in some experiments, the correct side won less than 50% of the time.

The study highlights how two single-agent measurements—claim wording threshold and message capacity—largely determine the collective fate in decision-making. By predicting collective transition points from single-agent measurements, it provides a new theoretical framework for analyzing collective behavior in language-model networks.

Nevertheless, the study also identifies limitations, such as prediction failures in some experiments and the inability of claim wording thresholds to transfer across different claim sets. Future research could explore larger-scale model behavior and the impact of different claim wordings on collective decisions.

Deep Analysis

Background

Recent advances in language models have significantly impacted natural language processing. However, ensuring model accuracy in collective decision-making remains challenging. Previous studies focused on individual model performance, overlooking the complexity of collective behavior.

Core Problem

In collective decision-making, language models can reach wrong consensus even when most agents initially are correct. The study explores how message capacity and claim wording affect this phenomenon.

Innovation

By combining message capacity and claim wording to predict collective truth-finding, the study provides a theoretical framework for analyzing collective behavior in language-model networks.

Methodology

  • �� Model message capacity to determine how many messages an agent reads.
  • �� Generate communication networks to simulate collective behavior.
  • �� Analyze the judgment process of an 8B parameter model using a stochastic binary neuron update rule.

Experiments

The experimental design includes 31,824 random queries to analyze the model's judgment as a logistic function. Additionally, 1,414 experiments verify prediction accuracy.

Results

Results indicate wrong consensus becomes unreachable when agents read fewer than 6.4 messages. However, in some experiments, the correct side won less than 50% of the time.

Applications

Findings can optimize language model applications in collective decision-making, especially in scenarios requiring high accuracy.

Limitations & Outlook

In some experiments, predictions failed with the correct side winning less than 50%. Claim wording thresholds failed to transfer across different claim sets.

Plain Language Accessible to non-experts

Imagine a group discussing an issue where each person can only hear part of the conversation. Even if most start correctly, the final decision might be wrong. The study finds that message capacity and claim wording are key factors. Message capacity determines how much each person can hear, while claim wording influences their judgment. It's like being in a crowded room where you can only hear a few voices, affecting your final judgment.

ELI14 Explained like you're 14

Imagine you're discussing a problem with friends, but you can only hear part of the conversation. Even if most people start correctly, the final decision might be wrong. The study finds that message capacity and claim wording are key factors. Message capacity determines how much each person can hear, while claim wording influences their judgment. It's like being in a crowded room where you can only hear a few voices, affecting your final judgment.

Glossary

Message Capacity

The number of messages an agent can read in a discussion.

Used to model how many messages an agent reads.

Claim Wording

The phrasing of a claim.

Affects an agent's initial judgment before reading messages.

Wrong Consensus

A collective decision that is incorrect.

Occurs even when most agents initially are correct.

Stochastic Binary Neuron

A model used to simulate an agent's judgment process.

Used to analyze the judgment process of an 8B parameter model.

Transition Point

The critical point where collective decision shifts from correct to incorrect.

Determined by message capacity and claim wording.

Open Questions Unanswered questions from this research

  • 1 How can this theoretical framework be applied to larger-scale models?
  • 2 How does claim wording affect collective decisions across different language models?

Applications

Immediate Applications

Collective Decision Optimization

Can be used to improve accuracy in language model collective decision-making.

Long-term Vision

Large-Scale Model Application

Explore the potential of applying this theory to larger-scale models.

Abstract

Whether human or large language model (LLM), an agent in a discussion reads only a few of the others' contributions, bounded by cognition, context, or cost. LLM collectives can settle on a wrong consensus even when a majority starts out correct; we ask how far that reading bound alone decides the outcome. We model the bound with one number, the message capacity, which sets how many of the others' messages an agent reads, and generate the communication network from it. Over 31,824 randomized queries, we found that an 8-billion-parameter model's judgment of a claim effectively reduces to a logistic function of a weighted sum of its inbox, the update rule of a stochastic binary neuron with divisively normalized weights. From these weights and the network's degree statistics alone, the wrong consensus should become unreachable from any start once agents read, on average, fewer than 6.4 of their 31 sources. In 1,414 episodes with assigned starts the prediction failed: the correct side won in fewer than 50% of episodes from every start, and in only 28-45% when 75% of agents started correct. The failure traces to the field, the threshold that a claim's wording sets for the agent's answer before any message is read: the experimental claims' fields lay below the calibration mean, and with each claim's own field the same weights reproduce the outcomes. Reversing the wording showed that the threshold follows what a claim asserts, not whether it is true. On a second 8B model the pipeline predicts claim-dependent bistability; transition points appeared where computed, and an eight-claim calibration matched in 15 of 16 conditions. At 70B the assertion bias is not detected. Thus a collective's fate is largely set by two single-agent measurements: the threshold a claim's wording sets, and the message capacity that sets the transition point.

cs.MA cs.CL nlin.AO physics.soc-ph