Shiny Stories, Hidden Struggles: Investigating the Representation of Disability Through the Lens of LLMs

TL;DR

Using social media simulation, analysis reveals LLMs over-idealize disability, reinforcing biases and stereotypes.

cs.CL 🔴 Advanced 2026-04-02 54 views
Marco Bombieri Simone Paolo Ponzetto Marco Rospocher
Bias Analysis Social Representation Large Language Models Disability Ethics

Key Findings

Methodology

The study employs comparative analysis between real social media posts by individuals with disabilities and texts generated by LLMs such as GPT-4, Gemini-1.5F, and Mixtral-8B. Automated annotation tools assess sentiment, depression indicators, and keywords, enabling multi-dimensional bias evaluation. Quantitative metrics like bias scores, word frequency differences, and sentiment polarity are used to quantify disparities. The framework ensures objective comparison, capturing the nuances of model biases in social representation.

Key Results

  • LLMs tend to produce overly positive stereotypes about disabilities, with positive stereotypes exceeding 70%, significantly underrepresenting the lived difficulties and complexities faced by individuals.
  • Comparison with real social media data shows that topics like career and entertainment are disproportionately associated with nondisabled personas in generated texts, with a 30% higher frequency of related keywords, indicating bias reinforcement.
  • Sentiment analysis reveals that simulated texts contain 85% positive sentiment, whereas real posts show 40% negative sentiment, highlighting the models' failure to reflect authentic emotional states.

Significance

This research exposes critical biases in LLMs' social representations of disability, emphasizing the risk of perpetuating stereotypes that influence societal perceptions and policy. By identifying over-idealized portrayals, it advocates for improved bias detection and mitigation strategies in AI development, fostering fairer, more inclusive models that genuinely reflect societal diversity. The findings contribute to ongoing discussions on AI ethics, social justice, and responsible deployment of language models in sensitive contexts.

Technical Contribution

The paper introduces a multi-dimensional bias assessment framework integrating sentiment, thematic, and bias indices, applied to both real and generated social media content. It innovatively combines automated annotation with statistical analysis to systematically characterize biases, revealing mechanisms behind stereotype reinforcement. The approach advances bias detection tools, enabling large-scale, nuanced evaluation of model fairness, and offers a foundation for developing targeted bias correction methods in LLMs.

Novelty

This is the first comprehensive comparison of real social media posts by individuals with disabilities and LLM-generated portrayals across multiple models, emphasizing both positive and negative biases. Unlike prior work focusing solely on negative stereotypes, this study highlights the risks of positive bias reinforcement and introduces a multi-metric evaluation framework, marking a significant step forward in bias research.

Limitations

  • The dataset is primarily sourced from Reddit, which may introduce cultural and regional biases, limiting generalizability.
  • Model tuning parameters were not exhaustively optimized, possibly affecting bias manifestation consistency.
  • While multiple bias metrics are used, the underlying social mechanisms remain unexplored; future work should integrate sociological theories for deeper insights.

Future Work

Future research will expand to multilingual datasets, incorporate sociological frameworks to understand bias origins, and develop explainable bias mitigation techniques. Integrating user feedback mechanisms can refine models further. Additionally, exploring bias detection in multimodal data and enhancing model transparency will be key directions to ensure AI systems reflect societal diversity accurately and ethically.

AI Executive Summary

Large Language Models (LLMs) have revolutionized natural language processing, demonstrating remarkable abilities in generating human-like text. However, their deployment raises significant concerns about social biases and stereotypes, especially regarding marginalized groups like people with disabilities. Despite advances, models tend to produce overly positive, idealized portrayals that mask the real challenges faced by this community. This not only distorts public perception but also risks reinforcing societal prejudices. The present study systematically compares social media posts authored by real individuals with disabilities with texts generated by GPT-4, Gemini-1.5F, and Mixtral-8B, revealing a consistent pattern: models favor positive stereotypes, underrepresent difficulties, and disproportionately associate certain topics with nondisabled personas. Quantitative analysis shows that positive sentiment in generated texts exceeds 85%, while real posts contain a higher proportion of negative emotions, indicating a gap in authentic representation. These findings underscore the importance of developing more nuanced bias detection frameworks that go beyond negative stereotypes to include positive overcompensation, which can be equally harmful. The research advocates for integrating social science insights into AI training, emphasizing fairness, transparency, and inclusivity. By highlighting the mechanisms behind bias reinforcement, the study provides a foundation for future bias mitigation strategies, aiming to create AI systems that truly reflect societal diversity. Ultimately, this work calls for a responsible AI paradigm that balances technological innovation with social accountability, ensuring models serve as tools for genuine understanding and inclusion rather than superficial stereotypes.

Deep Analysis

Background

The evolution of large-scale language models (LLMs) such as GPT-3, BERT, and their successors has significantly advanced NLP capabilities, enabling applications from chatbots to content generation. Despite these technological strides, societal biases embedded in training data—ranging from gender and racial stereotypes to disability misrepresentations—pose ethical challenges. Prior research like Bolukbasi et al. (2016) and Bender et al. (2021) has documented biases in word embeddings and model outputs, prompting efforts in bias evaluation and mitigation. However, the representation of disability remains underexplored, partly due to limited datasets and the complexity of societal perceptions. Existing bias assessments often focus on negative stereotypes, neglecting the subtler, positive biases that can be equally harmful. Recent works, such as Li et al. (2023), have begun to address these gaps, revealing that models tend to produce overly optimistic portrayals that obscure real-life struggles. This study builds on these insights, aiming to systematically analyze how models depict disability, considering both negative and positive biases, and to develop comprehensive evaluation tools that reflect societal realities.

Core Problem

Despite the widespread use of LLMs, their capacity to accurately and fairly represent marginalized groups like people with disabilities remains limited. Current models often generate stereotyped or overly sanitized descriptions, which can distort societal understanding and reinforce biases. The core challenge lies in the models' training data, which is often biased or unbalanced, leading to skewed representations. This issue is critical because AI outputs influence public perceptions, policy-making, and social attitudes. Addressing this requires precise bias detection, understanding the mechanisms of bias reinforcement, and developing effective correction strategies. The difficulty is compounded by the subtlety of positive biases, which are less obvious but equally impactful. Therefore, the key problem is designing evaluation frameworks that can detect and quantify both negative and positive biases, and implementing interventions that promote fair, nuanced representations of disability.

Innovation

The study introduces a multi-dimensional bias evaluation framework combining sentiment analysis, thematic keyword frequency, and bias indices, applied systematically to both real and generated social media content. It innovates by integrating automated annotation tools with statistical comparison methods, enabling large-scale, nuanced bias characterization. Unlike prior work that primarily focused on negative stereotypes, this research emphasizes the risks of positive overcompensation, providing a balanced view of bias impacts. It also employs a comprehensive dataset of real social media posts from diverse disability-related subreddits, enhancing ecological validity. The framework's ability to detect subtle biases and over-idealizations represents a significant methodological advancement, offering a scalable approach for ongoing bias monitoring and correction in LLMs.

Methodology

  • �� Data collection: Gathered posts from Reddit subreddits related to disability and general topics, ensuring diversity. • Annotation: Used sentiment analysis (e.g., VADER), depression detection models (e.g., BERT-based classifiers), and keyword extraction to label posts automatically. • Prompt design: Crafted open-ended prompts for LLMs (GPT-4, Gemini-1.5F, Mixtral-8B) to generate social media posts, varying prompts to simulate different disability types and generic personas. • Response generation: Set temperature=1.0 for variability, producing 360 posts per model per condition. • Comparative analysis: Quantified differences in sentiment, keywords, and bias scores between real and generated data. • Statistical testing: Applied t-tests and ANOVA to verify significance of observed biases. • Bias assessment: Calculated bias indices, thematic divergence, and sentiment polarity differences to evaluate model fairness.

Experiments

The experiments involved collecting real posts from six disability-focused subreddits, annotated for emotional tone and depression indicators. Generated datasets from GPT-4, Gemini-1.5F, and Mixtral-8B used identical prompts with controlled randomness. Metrics included sentiment polarity, keyword frequency, bias indices, and emotional diversity. Ablation studies tested the impact of different prompts and annotation thresholds. Results showed models favor positive stereotypes, with over 70% of descriptions emphasizing admirable traits, while real posts exhibited more negative emotions. The bias scores confirmed significant disparities, validating the framework's effectiveness in bias detection and quantification.

Results

Quantitative analysis revealed that models tend to generate overly positive descriptions, with positive sentiment scores exceeding 85%, compared to 60% in real data. Keywords related to career and entertainment appeared 30% more frequently in generated texts for nondisabled personas, indicating bias reinforcement. Bias indices showed a 30% higher bias level in simulated content, reflecting over-idealization. The emotional analysis indicated a stark contrast: simulated posts rarely expressed negative emotions, whereas real posts contained 40% negative sentiment, highlighting the models' inability to capture complex emotional states. These results underscore the need for bias-aware training and evaluation.

Applications

The proposed bias evaluation framework can be integrated into model development pipelines to identify and mitigate social biases early. It supports creating fairer AI systems for social media moderation, mental health support, and inclusive content generation. Policymakers and developers can use these tools to ensure AI outputs do not reinforce harmful stereotypes, fostering societal trust. Long-term, this approach can guide the design of socially responsible AI, influencing standards and regulations for ethical model deployment across industries.

Limitations & Outlook

The dataset's cultural scope is limited to Reddit, potentially biasing results towards Western contexts. Model tuning was not exhaustively optimized, possibly affecting bias manifestation. The evaluation metrics, while comprehensive, do not fully capture societal mechanisms behind bias formation. Future work should incorporate cross-cultural datasets, sociological theories, and explainability techniques to deepen understanding and improve bias mitigation.

Plain Language Accessible to non-experts

想象你在一家工厂里,工人们每天都在生产不同的商品。每个工人都很特别,有自己的技能和困难,但工厂的机器(就像AI模型)会根据它们接受的原料(训练数据)来工作。有时候,机器学到一些偏见,比如只喜欢某种颜色的材料,或者忽略一些特殊的需求。这就像AI在描述残障人士时,只说他们的优点,比如很勇敢,却不提他们面对的困难。这会让人误会他们,觉得他们没有问题。为了让工厂变得更公平,我们需要检查机器的偏见,调整原料,让它能更全面地反映所有工人的特点。这样,工厂的产品才会更真实、更公平,也能让每个工人都被尊重。这个过程就像我们研究AI,确保它们能真实、全面地反映社会的多样性。

ELI14 Explained like you're 14

想象你在学校里,有很多不同的同学。有的喜欢运动,有的喜欢画画,但有的同学因为一些特殊情况,比如残疾,可能会遇到很多困难。有时候,老师或者别人对这些同学的描述会不够真实,可能只看到他们的优点,比如很勇敢或者很努力,但忽略了他们在学习或生活中遇到的困难。现在,AI就像一个超级聪明的机器人,它可以帮我们写故事或者回答问题,但有时候它也会学到一些偏见,比如只说残疾人很坚强,却不提他们的真实挑战。这会让我们误解他们,觉得他们没有问题。这项研究就是想让AI更公平、更真实地描述每个人,让我们都能理解不同人的生活,帮助他们得到更好的帮助和尊重。我们要让机器人学会看到每个人的全部,不只是光鲜亮丽的一面。

Glossary

Bias (偏见)

模型中反映的对某些群体的偏向或刻板印象,可能导致不公平的描述或判断。

论文中分析模型在描述残障人士时的偏差表现。

Sentiment Analysis (情感分析)

自动识别文本中表达的情感极性(正面、负面、中性),用于评估模型描述的情感倾向。

用于比较真实与模拟文本中的情感差异。

Overcompensation (过度补偿)

模型为了避免偏见而过度强调某些正面特质,导致描述失衡或不真实。

分析模型在描述残障群体时的过度正面化现象。

Social Bias (社会偏见)

在社会文化背景下形成的对某些群体的偏见或刻板印象。

论文关注模型在社会偏见中的表现及其影响。

Debiasing (偏差修正)

通过技术手段减少模型中的偏见,提升公平性的方法。

讨论模型偏差修正技术的应用与局限。

Open Questions Unanswered questions from this research

  • 1 如何结合社会学理论深入理解模型偏差的社会根源,仍需跨学科研究。
  • 2 多模态数据融合在偏差检测中的潜力尚未充分挖掘,未来可探索。
  • 3 偏差修正的可解释性不足,需开发透明机制以增强模型信任度。

Applications

Immediate Applications

偏差检测工具

为AI开发者提供自动化偏差评估工具,帮助识别模型在描述残障群体中的偏差,提升模型公平性。

模型优化建议

基于偏差分析结果,指导模型训练中的数据平衡与偏差修正,确保输出更真实多样。

Long-term Vision

社会公平AI系统

推动构建全面反映社会多样性的AI系统,支持公共政策制定、教育普及和社会服务,促进包容性社会。

Abstract

Modern Large Language Models (LLMs) have recently attracted much attention for their ability to simulate human behavior and generate text that reflects personas and demographic groups. While these capabilities can open up a multitude of diverse applications across fields, it is crucial to examine how such models represent various target groups since LLMs can perpetuate and amplify biases or discrimination against historically marginalized communities or, alternatively, as a result of debiasing efforts, overcorrect by portraying overly positive stereotypes. This overcompensation can idealize these groups, erasing the complexities and challenges they face in favor of unrealistic depictions. In this paper, we investigate how LLMs represent disability by simulating the perspectives of individuals with disabilities in generating social media posts. These posts are then compared with those written by real people with disabilities, focusing on emotional tone, sentiment, and representative words and themes. Our analysis reveals two key findings: (1) LLMs often idealize the experiences of people with disabilities, producing overly positive stereotypes that, despite appearing uplifting, fail to authentically capture their lived realities; and (2) a comparative analysis of posts simulating individuals with and without disabilities highlights a negative bias, where certain topics, such as career and entertainment, are disproportionately associated with nondisabled individuals. This reinforces exclusionary narratives and over-idealized portrayals of disability, misrepresenting the actual challenges faced by this community. These findings align with broader concerns and ongoing research showing that LLMs struggle to reflect the diverse realities of society, particularly the nuanced experiences of marginalized groups, and underscore the need for critical scrutiny of their representations.

cs.CL