Cognitive models can reveal interpretable value trade-offs in language models

TL;DR

Cognitive models reveal value trade-offs in language models using RSA for polite speech analysis.

cs.CL 🔴 Advanced 2025-06-26 3 views
Sonia K. Murthy Rosie Zhao Jennifer Hu Sham Kakade Markus Wulfmeier Peng Qian Tomer Ullman
cognitive models language models value trade-offs polite speech interpretability

Key Findings

Methodology

The study employs Rational Speech Acts (RSA) models to analyze value trade-offs in language models. By manipulating reasoning budgets and system prompts, it evaluates models' performance in polite speech. An inverse reinforcement learning perspective is used to analyze training dynamics.

Key Results

  • Result 1: Models show higher informativeness (ϕ values) with increased reasoning budgets.
  • Result 2: Significant utility shifts occur early in training, with base model and pretraining data having a larger impact.
  • Result 3: Polite speech models can diagnose social behaviors like sycophancy.

Significance

This research provides new tools for understanding value trade-offs in language models, particularly in polite speech. The application of cognitive models offers better insights into controlling value trade-offs during training, guiding future model development.

Technical Contribution

The technical contribution lies in applying RSA models to analyze value trade-offs in language models, offering a new interpretative framework. The study reveals the impact of reasoning budgets and prompt manipulations on model behavior, providing new perspectives on value alignment.

Novelty

This study is the first to apply RSA models to analyze value trade-offs in language models, offering a systematic method to evaluate social behaviors in models.

Limitations

  • Limitation 1: Models may not fully capture human value trade-offs in complex social scenarios.
  • Limitation 2: The study focuses mainly on polite speech; applicability to other domains needs further validation.

Future Work

Future research could extend to other social behavior models, exploring value trade-offs across different domains. Further optimization of reasoning capabilities could enhance performance in complex scenarios.

AI Executive Summary

This study explores how cognitive models can reveal value trade-offs in language models, particularly in polite speech. Existing language models struggle with human-like value trade-offs, and the Rational Speech Acts (RSA) model from cognitive science offers a new analytical tool. By manipulating reasoning budgets and system prompts, the study evaluates different models' performance in polite speech. Results show that increased reasoning budgets lead models to favor informativeness, while early training utility shifts are mainly influenced by base models and pretraining data. This framework not only diagnoses social behaviors like sycophancy but also guides future model development.

The significance of this research lies in providing new tools and perspectives for value alignment in language models. By applying cognitive models, the study reveals the impact of reasoning budgets and prompt manipulations on model behavior, offering new insights into value alignment. The technical contribution is in applying RSA models to analyze value trade-offs, providing a new interpretative framework.

Despite significant achievements in polite speech, models may not fully capture human value trade-offs in complex social scenarios. Future research could extend to other social behavior models, exploring value trade-offs across different domains. Further optimization of reasoning capabilities could enhance performance in complex scenarios. This research provides new tools for understanding value trade-offs in language models, particularly in polite speech. By applying cognitive models, better insights into controlling value trade-offs during training are offered, guiding future model development.

Deep Analysis

Background

Language models struggle with human-like value trade-offs. The Rational Speech Acts (RSA) model from cognitive science offers a new analytical tool to explain human value trade-offs in language use. This study analyzes language models' performance in polite speech, revealing dynamic changes in value trade-offs.

Core Problem

Existing language models find it challenging to capture complex human value trade-offs, especially in polite speech. How cognitive models can reveal value trade-offs in language models is a critical research question.

Innovation

The study first applies RSA models to analyze value trade-offs in language models, offering a systematic method to evaluate social behaviors. By manipulating reasoning budgets and system prompts, it reveals models' performance in polite speech.

Methodology

  • �� Use RSA models to analyze value trade-offs in language models.
  • �� Manipulate reasoning budgets and system prompts to evaluate models' performance in polite speech.
  • �� Use an inverse reinforcement learning perspective to analyze training dynamics.

Experiments

The experimental design includes manipulating reasoning budgets and system prompts in both closed and open-source models, evaluating their performance in polite speech. An inverse reinforcement learning perspective is used to analyze training dynamics.

Results

Results show that increased reasoning budgets lead models to favor informativeness, while early training utility shifts are mainly influenced by base models and pretraining data. Polite speech models can diagnose social behaviors like sycophancy.

Applications

The research provides new tools and perspectives for value alignment in language models, particularly in polite speech. It can be used to develop more socially sensitive language models.

Limitations & Outlook

Despite significant achievements in polite speech, models may not fully capture human value trade-offs in complex social scenarios.

Plain Language Accessible to non-experts

Imagine you're in a kitchen cooking. You need to balance healthiness and taste. Cognitive models are like a chef helping language models balance informativeness and social utility. By adjusting ingredients (reasoning budgets and prompts), the chef changes the dish's flavor (model behavior). It's like making a dish that's both healthy and tasty; models need to find the best balance between providing accurate information and maintaining social politeness.

ELI14 Explained like you're 14

Imagine you're playing a game where you have to choose between scoring points and making friends. Cognitive models are like your game guide, helping you balance the two. By adjusting game strategies (reasoning budgets and prompts), you change the game's outcome (model behavior). It's like winning the game while keeping your friends; models need to find the best balance between providing accurate information and maintaining social politeness.

Glossary

Rational Speech Acts (RSA) Model

A cognitive model used to explain value trade-offs in human language use.

Used to analyze language models' performance in polite speech.

Reasoning Budget

The amount of computational resources available to a model during decision-making.

Used to evaluate model behavior under different reasoning budgets.

System Prompt

Instructions or suggestions used to guide model behavior.

Manipulated to evaluate value trade-offs in models.

Inverse Reinforcement Learning

A technique to infer goals or utilities by observing behavior.

Used to analyze training dynamics in models.

Sycophancy

Behavior characterized by excessive flattery or servility.

Diagnosed through polite speech models.

Open Questions Unanswered questions from this research

  • 1 How can models more accurately capture human value trade-offs in complex social scenarios?
  • 2 How do existing models perform in other social behaviors?
  • 3 How to optimize models' reasoning abilities to improve performance in complex scenarios?

Applications

Immediate Applications

Social Media Assistant

Can be used to develop more socially sensitive social media assistants, enhancing user experience.

Long-term Vision

Human-Machine Interaction Optimization

Optimize human-machine interaction experiences by better understanding human value trade-offs.

Abstract

Value trade-offs are an integral part of human decision-making and language use, however, current tools for interpreting such dynamic and multi-faceted notions of values in language models are limited. In cognitive science, so-called "cognitive models" provide formal accounts of such trade-offs in humans, by modeling the weighting of a speaker's competing utility functions in choosing an action or utterance. Here, we show that a leading cognitive model of polite speech can be used to systematically evaluate alignment-relevant trade-offs in language models via two encompassing settings: degrees of reasoning "effort" and system prompt manipulations in closed-source frontier models, and RL post-training dynamics of open-source models. Our results show that LLMs' behavioral profiles under the cognitive model a) shift predictably when they are prompted to prioritize certain goals, b) are amplified by a small reasoning budget, and c) can be used to diagnose other social behaviors such as sycophancy. Our findings from LLMs' post-training dynamics reveal large shifts in values early on in training and persistent effects of the choice of base model and pretraining data, compared to feedback dataset or alignment method. Our framework offers a flexible tool for probing behavioral profiles across diverse model types and gaining insights for shaping training regimes that better control trade-offs between values during model development.

cs.CL cs.AI