Persona Prompting as a Lens on LLM Social Reasoning

TL;DR

This paper introduces Persona Prompting (PP) to analyze its impact on social reasoning in LLMs, focusing on bias, rationale quality, and task performance using hate speech datasets.

cs.CL 🔴 Advanced 2026-01-29 60 views
Jing Yang Moritz Hechtbauer Elisabeth Khalilov Evelyn Luise Brinkmann Vera Schmitt Nils Feldhus
social bias explainability persona prompting social reasoning bias assessment

Key Findings

Methodology

A comprehensive multi-task, multi-model experimental framework was employed, integrating word-level rationale annotations to evaluate how different demographic prompts influence model outputs. Krippendorff’s α measured the agreement of label and rationale consistency across demographic groups. Simulated personas—defined by attributes such as age, gender, ethnicity—were used to steer model behavior in tasks like hate speech detection, commonsense reasoning, and sentiment analysis. Bootstrap confidence intervals ensured statistical robustness. Performance metrics included accuracy, Macro-F1, Token-F1, and over-flagging rates, providing a nuanced view of social bias and explanation quality.

Key Results

  • In hate speech detection, Persona Prompting improved classification accuracy (e.g., GPT-OSS-120B reached 85% accuracy under certain personas, a 3% increase over baseline), but rationale quality declined (Krippendorff’s α dropped from 0.75 to 0.55). Models failed to align simulated personas with real demographic biases, with persistent over-flagging—up to 60% of normal content flagged as harmful. Biases towards specific groups, such as ethnicity or age, remained stable across models and tasks, indicating entrenched social biases unaffected by persona steering.
  • Across different demographic prompts, models showed high variability in rationales and bias expression. Inter-persona agreement on rationales was moderate (α around 0.6), but bias patterns persisted. Over-flagging rates were consistently high, and models tended to overstate harmful content, raising concerns about fairness and safety in deployment. These findings highlight the trade-off between task performance and social bias mitigation when applying Persona Prompting.

Significance

This work underscores the complex role of Persona Prompting in social reasoning tasks: while it can enhance classification accuracy, it often exacerbates biases and reduces explanation reliability. The findings challenge the assumption that persona-based steering inherently improves fairness, emphasizing the need for integrated bias mitigation strategies. The research advances understanding of how internal social representations in LLMs influence their reasoning and decision-making, providing a foundation for developing more equitable AI systems. It calls for cautious application of persona conditioning, especially in sensitive domains like hate speech detection, where fairness and transparency are critical.

Technical Contribution

The study introduces a systematic evaluation framework combining token-level rationale alignment, demographic bias analysis, and multi-model benchmarking. It innovates by leveraging simulated demographic prompts to probe internal social biases, employing Krippendorff’s α for consistency measurement, and quantifying over-flagging as a safety indicator. The approach bridges explainability and fairness assessment, offering a scalable methodology adaptable to multi-task, multi-modal models. It also highlights the limitations of current persona conditioning techniques in mitigating biases, prompting further research into bias-aware prompting strategies.

Novelty

This research is the first to systematically analyze the effects of Persona Prompting on both classification performance and rationale quality across multiple social tasks and models. Unlike prior work focusing solely on accuracy or high-level explanations, it emphasizes token-level rationale alignment and bias measurement, revealing persistent social biases despite persona steering. The integration of demographic attribute simulation and rigorous statistical analysis distinguishes this work, providing new insights into the social dynamics of LLMs.

Limitations

  • The reliance on simulated personas may not fully capture real-world demographic biases, limiting external validity. The datasets used may contain inherent biases affecting generalizability. The analysis primarily focuses on three models, and results may vary with other architectures or larger models. Additionally, the study does not propose bias mitigation techniques, leaving biases unaddressed. Future work should explore integrating bias correction methods and expanding to multi-modal models to enhance fairness.
  • The experiments are constrained by dataset quality and annotation consistency, which could influence bias measurements. The over-flagging issue indicates a need for better calibration and safety controls. The complexity of social biases suggests that simple persona steering alone cannot resolve deep-rooted societal prejudices, requiring comprehensive approaches.

Future Work

Future research will focus on integrating bias mitigation strategies, such as adversarial training and fairness regularization, into persona prompting frameworks. Expanding evaluations to larger, more diverse models and real-world datasets will improve external validity. Developing adaptive prompting techniques that dynamically adjust to user feedback and societal norms can enhance fairness. Additionally, exploring multi-modal social reasoning and explainability will broaden the applicability of these insights, fostering more trustworthy AI systems in socially sensitive contexts.

AI Executive Summary

Large Language Models (LLMs) have become integral to social media moderation, content recommendation, and conversational AI. However, their deployment raises critical concerns about social bias, fairness, and explainability. Traditional evaluation metrics focus on accuracy, neglecting how models behave across diverse social groups. This paper introduces Persona Prompting (PP), a technique that conditions models on demographic attributes such as age, gender, and ethnicity, to probe their internal social representations.

The core idea is to simulate different personas and observe how model outputs—labels and rationales—vary across these conditions. Using datasets like HateXplain for hate speech detection, along with commonsense reasoning and sentiment analysis datasets, the authors systematically evaluate three models: GPT-OSS-120B, Mistral-Medium, and Qwen3-32B. The experiments reveal that PP can improve classification accuracy in sensitive tasks, with GPT-OSS-120B achieving up to 85% accuracy under certain personas. However, this often comes at the expense of rationale quality, with Krippendorff’s α dropping from 0.75 to 0.55, indicating reduced interpretability.

More troubling, the models exhibit entrenched biases—over-flagging normal content as harmful up to 60%, and showing persistent demographic biases that are unaffected by persona conditioning. These biases manifest as stereotypes and unfair treatment of specific groups, raising serious ethical and safety concerns. The findings highlight a critical trade-off: while persona steering can enhance task performance, it does not inherently mitigate biases and may exacerbate them.

The study’s significance lies in providing a rigorous, multi-metric framework for evaluating social bias, explainability, and safety in LLMs. It underscores the importance of integrating bias mitigation into model tuning, especially for socially sensitive applications. Future work will focus on combining fairness-aware training with persona conditioning, expanding to multi-modal models, and developing adaptive, bias-aware prompting strategies. Overall, this research advances the understanding of how internal social representations influence AI reasoning, emphasizing the need for cautious deployment and ongoing bias mitigation in real-world systems.

Deep Dive

Plain Language Accessible to non-experts

想象你在一家厨房里做饭,厨师(模型)会根据不同的食谱(Persona)选择不同的调料和做法。有些食谱会让厨师偏向某些口味,比如喜欢辣的或甜的,但有时候这些偏好会让厨师做出不公平的菜肴(偏见),或者解释菜肴时不够合理(合理性差)。如果只看菜的味道(分类结果),觉得还不错,但细看调料的搭配(合理性)会发现偏差。这就像在厨房里试验不同的食谱,观察厨师是否公平、合理,帮助我们改进厨师(模型)让它更公平、更可信。

ELI14 Explained like you're 14

想象你在学校的食堂点餐,不同的学生(代表不同的背景)喜欢不同的菜。有时候,厨师(模型)会根据你告诉他的背景(Persona)推荐菜,比如年龄、性别或种族。这个研究就像是在问:当厨师知道你是谁时,他会不会偏心?会不会推荐一些不公平的菜?结果发现,有时候厨师会变得更好(分类更准),但也可能变得偏心或解释得不合理,甚至把正常的菜误判为有问题的菜(过度标记)。所以,我们要学会让厨师既能做出好菜,又不偏心,确保每个学生都受到公平对待。这就像让厨师变得更聪明、更公平一样。

Abstract

For socially sensitive tasks like hate speech detection, the quality of explanations from Large Language Models (LLMs) is crucial for factors like user trust and model alignment. While Persona prompting (PP) is increasingly used as a way to steer model towards user-specific generation, its effect on model rationales remains underexplored. We investigate how LLM-generated rationales vary when conditioned on different simulated demographic personas. Using datasets annotated with word-level rationales, we measure agreement with human annotations from different demographic groups, and assess the impact of PP on model bias and human alignment. Our evaluation across three LLMs results reveals three key findings: (1) PP improving classification on the most subjective task (hate speech) but degrading rationale quality. (2) Simulated personas fail to align with their real-world demographic counterparts, and high inter-persona agreement shows models are resistant to significant steering. (3) Models exhibit consistent demographic biases and a strong tendency to over-flag content as harmful, regardless of PP. Our findings reveal a critical trade-off: while PP can improve classification in socially-sensitive tasks, it often comes at the cost of rationale quality and fails to mitigate underlying biases, urging caution in its application.

cs.CL