Auditing Support Strategies in LLMs through Grounded Multi-Turn Social Simulation
Proposes grounded multi-turn social simulation with SSBC to analyze LLM support strategy shifts based on estimated user distress.
Key Findings
Methodology
The framework decomposes Reddit support narratives into ordered fragments, simulating multi-turn responses from LLMs (Llama-3.1-8B, OLMo-3-7B). It employs SSBC for multi-label behavior classification and linear probes to estimate internal distress signals from hidden states. Over 6200 turns, the study analyzes how support strategies change with estimated distress, revealing significant declines in teaching (−27.4%) and increases in validation (+31.0%) and empathy (+24.8%) as distress rises. Community context influences behavior independently of demographics, emphasizing the importance of trajectory-level analysis over single-turn evaluation.
Key Results
- Support strategies systematically shift with estimated distress: teaching declines sharply (−27.4 pp), while validation and emotional support increase significantly, indicating a behavioral re-weighting under stress.
- Model architecture consistency confirms robustness; community norms, rather than demographics, shape behavior preferences, with notable differences in advice and encouragement across subreddits.
- Trajectory analysis uncovers dynamic support patterns invisible in single-turn assessments, highlighting the necessity of multi-turn frameworks for social support evaluation.
Significance
This work advances the understanding of multi-turn support behavior in LLMs, addressing limitations of static single-response evaluations. By quantifying how support strategies adapt during ongoing interactions, it provides a foundation for safer, more reliable AI in sensitive domains like mental health. The internal distress estimation approach offers a novel way to interpret model behavior, guiding future improvements in alignment and safety. Overall, it bridges a critical gap between theoretical social support models and practical AI deployment, fostering trust and efficacy in AI-assisted social interventions.
Technical Contribution
The study introduces a grounded simulation framework integrating real narratives, multi-label behavior coding, and internal signal estimation via linear probes. It systematically quantifies support strategy dynamics across multiple models and communities, establishing a new paradigm for trajectory-level auditing. The combination of real data grounding and internal state analysis distinguishes this approach from prior static or synthetic evaluation methods, enabling nuanced insights into model behavior under realistic, multi-turn scenarios.
Novelty
This is the first application of grounded multi-turn social simulation combined with SSBC and internal distress probes to audit LLM support behavior. Unlike previous single-turn or synthetic dialogue evaluations, it emphasizes trajectory analysis, revealing behavioral shifts aligned with user distress levels. The integration of real Reddit narratives and multi-label coding offers a comprehensive, scalable framework for dynamic behavior assessment, marking a significant innovation in AI social support research.
Limitations
- The internal distress estimation relies on linear probes, which may not fully capture complex internal representations, potentially limiting accuracy under certain contexts.
- Community effects are inferred from discourse norms, not causal demographic analysis, leaving room for confounding factors.
- Experiments are limited to mid-scale models; larger models or multi-task settings may exhibit different behavior patterns, requiring further validation.
Future Work
Future research will incorporate multimodal signals, larger models, and personalized adaptation mechanisms. Exploring real-time feedback integration and extending to diverse social domains will enhance robustness. Additionally, developing causal analysis methods to disentangle community and demographic influences will refine understanding of bias and fairness in social support models.
AI Executive Summary
In recent years, large language models (LLMs) have shown promise in providing social support through conversational agents. However, traditional evaluation methods focus on single-turn responses, neglecting the dynamic nature of real-world interactions where users gradually disclose their issues. This gap limits understanding of how models adapt their support strategies over multiple turns, especially under varying levels of user distress. To address this, the authors propose a grounded multi-turn social simulation framework that decomposes real Reddit support narratives into ordered fragments, simulating multi-turn interactions with LLMs such as Llama-3.1-8B and OLMo-3-7B.
The core innovation lies in combining SSBC, a multi-label taxonomy of support behaviors, with linear probes that estimate the model’s internal perception of user distress from hidden states. This approach enables a trajectory-level analysis of how support strategies shift as perceived distress increases. Empirical results reveal a systematic decline in instructional support (teaching) by 27.4 percentage points, while emotional and validation strategies increase significantly, indicating a behavioral re-weighting under stress. These patterns are consistent across models and are influenced more by community norms than demographic factors.
This research highlights the importance of multi-turn evaluation for socially sensitive AI applications. It demonstrates that support behaviors evolve dynamically, which single-turn assessments fail to capture. The findings inform future model design and safety protocols, emphasizing the need for trajectory-aware auditing tools. Limitations include reliance on internal probes and scope restricted to mid-scale models. Future directions involve multimodal signals, larger architectures, and personalized support mechanisms, aiming to build safer, more adaptive AI social support systems that can better serve users in complex, real-world scenarios.
Deep Analysis
Background
随着大规模语言模型(LLMs)在对话系统中的广泛应用,支持性对话的评估逐渐成为研究焦点。早期工作多关注单轮响应质量,如Wang等(2025)和Lee等(2024)提出的指标,但忽视了多轮交互中支持策略的演变。近年来,社会支持理论强调逐步披露信息和动态调节策略的重要性(Cutrona & Russell, 1990),但缺乏系统的模型化和量化工具。现有社会模拟框架如Sotopia和SimulatorArena虽能模拟多轮交互,但多依赖于人工设计或生成式用户,难以真实反映用户行为复杂性。本文基于真实Reddit支持帖子的多轮模拟,结合SSBC多标签编码,提出一种 grounded social simulation 方法,旨在揭示模型在逐步披露信息中的行为轨迹,为模型调优提供理论基础。
Core Problem
当前大部分支持模型评估集中于单轮响应,无法捕捉多轮对话中支持策略的变化和潜在偏差。尤其在心理健康等敏感场景,模型可能在不同阶段表现出不同的行为偏向,影响用户体验和安全性。缺乏系统化的行为轨迹分析工具,难以识别模型在逐步披露信息时的策略偏差和潜在风险。如何建立一个既能反映真实用户行为,又能量化模型行为变化的评估框架,成为亟待解决的问题。这不仅关系到模型的可信度,也影响其在实际社会支持中的应用效果。
Innovation
本文创新点在于提出Grounded Multi-Turn Social Simulation框架,结合真实支持叙事和多标签行为编码,利用线性探针估算模型内部困扰感知,突破传统静态单轮评价的局限。具体创新包括:
- �� 真实叙事拆解:基于真实Reddit帖子,避免合成数据偏差。
- �� 多标签行为编码:采用SSBC,细粒度捕捉支持策略。
- �� 内部信号估算:利用线性探针分析模型困扰感知,反映潜在偏差。
- �� 轨迹分析:揭示支持策略随困扰变化的动态趋势,为模型调优提供依据。这些创新使得模型行为的动态演变得以量化和理解,推动多轮对话支持模型的研究前沿。
Methodology
- �� 数据准备:采集五个Reddit社区的支持帖,人工标注困扰程度,拆解为有序碎片。
- �� 多轮模拟:将碎片逐轮输入模型(Llama-3.1-8B、OLMo-3-7B),每轮生成响应。
- �� 行为编码:用SSBC对每轮响应进行多标签分类,识别支持策略类型。
- �� 内部信号估算:训练线性探针,从隐藏层提取表示,估算模型对用户困扰的内在感知。
- �� 统计分析:使用χ2检验和混合效应逻辑回归,分析支持策略与困扰的关系及社区差异。
- �� 轨迹分析:结合支持行为变化和社区语境,揭示动态趋势。
Experiments
采用五个Reddit社区的支持帖子,标注困扰程度,拆解为平均6.7个碎片。模型响应逐轮生成,行为由SSBC编码,训练线性探针估算困扰信号。对6200多轮数据进行统计检验,比较不同困扰水平下的支持策略变化。模型架构包括Llama-3.1-8B和OLMo-3-7B,评估指标为SSBC标签的变化率和统计显著性。通过社区差异分析,验证策略偏好受社区语境影响。实验还包括不同模型和参数设置的对比,确保结果的稳健性。
Results
支持策略随困扰升高显著变化,教学策略下降27.4个百分点,验证和情感策略上升超30个百分点,表明模型在高困扰时偏向安抚而非指导。社区语境影响明显,偏向话题和话语规范。模型在不同架构中表现一致,轨迹变化在单轮评估中难以察觉,强调多轮模拟的重要性。这些发现验证了模型行为的动态调节能力,为模型安全性和伦理性提供新思路。
Applications
该框架可应用于心理健康支持、在线辅导、情感分析等场景,帮助开发者识别模型在不同交互阶段的偏差,优化支持策略,确保模型在敏感场景中的安全性。未来可结合个性化调节和多模态信息,提升模型的适应性和可信度,推动智能社会支持系统的落地。
Limitations & Outlook
模型对困扰的估算依赖线性探针,可能受模型内部表示偏差影响,未考虑多模态信息。社区语境分析偏重话题和话语规范,未深入探讨人口统计特征影响。实验范围局限于中等规模模型,未来需扩展到更大模型和多任务场景,提升泛化能力。
Plain Language Accessible to non-experts
想象你在一家工厂工作,工厂里有很多不同的机器,每台机器都能做不同的事情。有时候,工厂需要根据不同的任务调整机器的工作方式,比如生产不同的产品。这个研究就像是在观察这些机器在不同任务中的表现变化。我们用一种方法,把工厂里工人的对话拆成一段段的故事,然后模拟机器(模型)逐步回应工人的请求。通过观察这些回应,我们发现当工人表现出更大的困扰时,机器会更倾向于安慰和鼓励,而不是教导或提供具体的解决方案。这个过程帮助我们理解,机器在不同情况下会采取不同的支持策略,就像工厂里的机器会根据任务调整工作方式一样。这种方法可以帮助我们确保未来的支持系统更贴合人们的真实需求,更安全、更有效。
ELI14 Explained like you're 14
想象你在学校里,有个老师会帮助你解决问题。有时候,你会告诉老师你很困惑,老师会用不同的方式帮你,比如安慰你、鼓励你,或者告诉你怎么做。这个研究就像是在观察老师在不同情况下会用什么方法帮助你。科学家们用一种特别的方法,把你和老师的对话拆成一段段,然后模拟老师逐步回应你。通过观察这些回应,他们发现当你表现得很困扰时,老师会更倾向于安慰和鼓励你,而不是马上告诉你具体怎么做。这样一来,老师会根据你的情绪变化,调整帮助你的方式。这就像是让老师变得更聪明,知道什么时候该安慰,什么时候该教你技能。这个研究帮助我们理解,未来的聊天机器人也可以像老师一样,根据你的心情,提供更贴心的帮助,让你感觉更安心、更被理解。
Glossary
Social Support Behavior Code (SSBC) (社会支持行为编码)
一种多标签分类体系,用于细粒度识别支持对话中的不同策略,如安慰、验证、教学等。它帮助分析模型在多轮对话中的行为组成。
在论文中,SSBC被用来编码每轮模型响应的支持策略类型。
线性探针 (Linear Probe)
一种分析工具,通过训练线性分类器,从模型隐藏层的表示中估算特定信号(如困扰感知),不影响模型生成过程。
用于估算模型对用户困扰的内部感知信号,分析行为变化。
Grounded Multi-Turn Social Simulation (基于真实叙事的多轮社会模拟)
一种模拟框架,将真实支持帖子的叙事拆解为有序碎片,逐轮模拟模型响应,揭示行为轨迹。
核心创新,用于分析模型在逐步披露信息中的行为动态。
支持策略 (Support Strategies)
在对话中采取的不同行为,如提供信息、情感安慰、验证、教学等,用于满足用户不同需求。
通过SSBC识别模型在多轮交互中的支持策略变化。
Open Questions Unanswered questions from this research
- 1 未来研究需结合多模态信息,提升模型对复杂情境的理解能力,特别是在情感、语调等多模态信号的融合方面。
Applications
Immediate Applications
心理健康支持系统
利用多轮模拟框架,优化聊天机器人在心理咨询中的行为策略,确保在不同阶段提供适当的情感支持和指导,提升用户体验。
在线辅导平台
通过行为轨迹分析,检测模型在多轮交互中的偏差,及时调整支持策略,保障服务安全和效果。
Long-term Vision
智能社会支持网络
结合多模态信息和个性化调节,构建全场景、多轮支持的智能系统,全面提升社会援助的效率和公平性。
Abstract
When users seek social support from chatbots, they disclose their situation gradually, yet most evaluations of supportive LLMs rely on single-turn, fully specified prompts. We introduce a multi-turn simulation framework that closes this gap. Support-seeking narratives from five Reddit communities are decomposed into ordered fragments and revealed turn by turn to a language model. Each response is coded with the Social Support Behavior Code (SSBC), an established multi-label taxonomy that captures the composition of support, rather than a single quality score. To ask whether support choices track the model's own construal of user distress, we use linear probes on hidden representations to estimate this internal signal without altering the generation context. Across two mid-scale models (Llama-3.1-8B, OLMo-3-7B) and more than 6,200 turns, support composition shifts systematically with estimated distress: teaching declines as estimated distress rises, a finding that replicates across architectures, while increases in affective and esteem-oriented strategies (such as validation) are suggestive but model-specific and rest on noisier annotations. Community context independently shapes behavior, tracking topic and discourse norms rather than demographic categories. These trajectory-level dynamics, invisible to single-turn evaluation, motivate multi-turn auditing frameworks for socially sensitive applications.