Epistemic Stance Flexibility Probing: Measuring Prompt-Conditioned Register Shift in Large Language Models
ESFP benchmark measures epistemic stance flexibility in large language models under different prompts.
Key Findings
Methodology
The ESFP benchmark evaluates models by contrasting externally attributed and self-attributed prompts across four dimensions: lexical self-attribution, role framing responsiveness, sentence-level stance content density, and cross-condition stance consistency.
Key Results
- Among eight frontier models, a 27B open-weight model matches the strongest proprietary systems, indicating epistemic flexibility is largely orthogonal to model capability.
- Stance content density provides the strongest signal, while surface-level lexical markers like 'I think' change significantly without altering expressed stance.
- Item-level bootstrap confidence intervals and weight-sensitivity analyses are provided.
Significance
ESFP benchmark measures a model's propensity to adapt its epistemic stance under changing attribution conditions, rather than a general competence measure. Higher scores should not be interpreted as universally better.
Technical Contribution
Introduces a new behavioral benchmark capable of evaluating model behavior by varying the frame while holding content constant.
Novelty
First to measure language model epistemic stance flexibility through prompt-conditioned changes.
Limitations
- Models perform poorly on consistency, potentially due to noise rather than anti-consistency.
- Limited to English lexicon, potentially affecting non-English model evaluations.
Future Work
Future work could expand to multilingual settings and incorporate human gold labels to enhance evaluation accuracy.
AI Executive Summary
In the context of modern AI, language models are required to exhibit different epistemic stances under varying prompts. However, existing benchmarks do not directly assess this capability.
The ESFP benchmark evaluates models by contrasting externally attributed and self-attributed prompts across four dimensions: lexical self-attribution, role framing responsiveness, sentence-level stance content density, and cross-condition stance consistency. The study finds that epistemic flexibility is largely orthogonal to model capability, with a 27B open-weight model matching the strongest proprietary systems.
This research provides a measure of a model's propensity to adapt its epistemic stance under changing attribution conditions, rather than a general competence measure. Future work could expand to multilingual settings and incorporate human gold labels to enhance evaluation accuracy.
Deep Analysis
Background
Language models need to exhibit different epistemic stances in conversations based on prompt conditions. Existing benchmarks focus on accuracy, instruction following, and safety, but do not directly assess this capability.
Core Problem
The core problem is whether models can exhibit consistent and reasonable changes in epistemic stance under different prompt conditions.
Innovation
The ESFP benchmark evaluates models by contrasting externally attributed and self-attributed prompts across four dimensions, achieving the first measurement of epistemic stance flexibility.
Methodology
- �� Lexical self-attribution: Analyzes the vocabulary used by models under different prompts.
- �� Role framing responsiveness: Evaluates model response to changes in role framing.
- �� Sentence-level stance content density: Assessed by an LLM judge panel.
- �� Cross-condition stance consistency: Evaluated using Fleiss’ κ.
Experiments
The experiment evaluated eight frontier models using 520 independent prompts, covering six epistemic categories and five phrasing templates.
Results
A 27B open-weight model matches the strongest proprietary systems, with stance content density providing the strongest signal.
Applications
Can be used to evaluate conversational agents' adaptability under different prompt conditions.
Limitations & Outlook
Limited to English lexicon, potentially affecting non-English model evaluations.
Plain Language Accessible to non-experts
Imagine a smart assistant that needs to show different attitudes in different situations. When asked about expert opinions, it should give neutral answers; when asked for its own opinion, it should express its views. The ESFP benchmark is like a measuring tool that helps us understand if these assistants can flexibly adjust their attitudes under different prompts.
ELI14 Explained like you're 14
Imagine you're playing a role-playing game where your character sometimes needs to give its own opinion and sometimes report others'. ESFP is like a game rule that helps us see if these characters can behave appropriately in different tasks. Just like you need to switch roles in different game tasks, ESFP tests models' performance under different prompts.
Glossary
ESFP (Epistemic Stance Flexibility Probing)
A benchmark to evaluate language models' changes in epistemic stance under different prompt conditions.
Used to measure model responses to externally attributed and self-attributed prompts.
Lexical Self-Attribution
Analyzes the vocabulary used by models to assess their tendency for self-attribution.
Used to evaluate vocabulary changes under different prompts.
Sentence-Level Stance Content Density
Assessed by a judge panel to measure the stance content at the sentence level.
Used to evaluate stance content changes under different prompts.
Cross-Condition Stance Consistency
Evaluated using Fleiss’ κ to ensure consistent stance changes across different prompts.
Ensures coherent stance changes under different prompts.
Fleiss’ κ
A statistical method to evaluate agreement among multiple raters.
Used to evaluate stance consistency across different prompt conditions.
Open Questions Unanswered questions from this research
- 1 How to evaluate epistemic stance flexibility in multilingual settings?
- 2 How to enhance evaluation accuracy, especially for non-English models?
Applications
Immediate Applications
Conversational Agent Evaluation
Used to evaluate conversational agents' adaptability under different prompt conditions, aiding in developing smarter dialogue systems.
Long-term Vision
Multilingual Model Evaluation
Expand to multilingual settings to evaluate epistemic stance flexibility across different language models.
Abstract
A language model may be asked either what experts believe about a contested claim or what it believes about the claim itself. A trustworthy conversational agent should distinguish these two requests and respond in different epistemic registers: neutral attribution in the first case and stance expression in the second. Whether such a shift occurs-and whether it occurs coherently-is not directly assessed by existing benchmarks for accuracy, instruction following, or safety. We introduce ESFP, a behavioral benchmark that treats the contrast between externally attributed and self-attributed prompts as the fundamental unit of measurement. ESFP consists of 104 carefully controlled items spanning six epistemic categories and five phrasing templates, and evaluates model responses along four complementary dimensions: lexical self-attribution, representation-level responsiveness to role framing, sentence-level stance content density assessed by an LLM judge panel, and cross-condition stance consistency. Evaluating eight frontier models from five vendors, we find that epistemic flexibility is largely orthogonal to general model capability: a 27B open-weight model matches the strongest proprietary systems, the flagship model of one family underperforms its lightweight counterpart, and reasoning-optimized models do not consistently exhibit higher flexibility. Stance content density provides the strongest signal, while surface-level lexical markers such as 'I think' can change substantially without corresponding changes in expressed stance. We provide item-level bootstrap confidence intervals, weight-sensitivity analyses, and an explicit discussion of the interpretation limits of the composite score. ESFP measures a model's propensity to adapt its epistemic stance under changing attribution conditions, rather than a general competence measure.