Algorithmic Fragility and Persona Bias in LLM-Generated Autistic Communication
Study reveals algorithmic fragility and persona bias in LLM-generated autistic communication using a multi-agent qualitative analysis framework.
Key Findings
Methodology
The study employs a dual-persona rewrite paradigm, prompting ten LLMs to rewrite autistic discourse under autistic and neurotypical personas. A multi-agent qualitative analysis framework analyzes 14,840 rewrite pairs, revealing persona-specific generative breakdowns.
Key Results
- Autistic persona rewrites diverge significantly in lexical form and affective register compared to neurotypical rewrites, despite equivalent semantic similarity.
- Most models collapse cross-persona generations into near-identical outputs, indicating persona-specific generative breakdown.
- Comparison with autistic human annotators shows community-insider knowledge leads to systematic label reversals relative to LLM classifications.
Significance
The study reveals that current alignment training causes persona-specific generative breakdowns visible only through qualitative analysis, confirming a deep representational gap that prompt engineering cannot resolve. This has significant implications for authentic autistic communication representation.
Technical Contribution
The study introduces a multi-agent qualitative analysis framework, uncovering systemic output erasure, stereotyped hallucination, and task-evasive meta-commentary as failure modes in LLM-generated autistic communication.
Novelty
This is the first study to reveal persona-specific breakdowns in LLM-generated autistic communication using a dual-persona rewrite paradigm, offering a novel analytical perspective.
Limitations
- Outputs generated under the autistic persona deviate in lexical and affective aspects, indicating persona-specific bias.
- Most models produce nearly identical outputs across personas, failing to capture persona differences.
Future Work
Future research could explore improved alignment strategies to reduce persona-specific generative breakdowns and develop more inclusive models.
AI Executive Summary
The study reveals algorithmic fragility and persona bias in LLM-generated autistic communication. Using a dual-persona rewrite paradigm, researchers prompted ten LLMs to rewrite autistic discourse under autistic and neurotypical personas. Results show that autistic persona rewrites diverge significantly in lexical form and affective register compared to neurotypical rewrites, despite equivalent semantic similarity.
The study introduces a multi-agent qualitative analysis framework, analyzing 14,840 rewrite pairs and uncovering systemic output erasure, stereotyped hallucination, and task-evasive meta-commentary as failure modes. Comparison with autistic human annotators shows that community-insider knowledge leads to systematic label reversals relative to LLM classifications.
The findings indicate that current alignment training causes persona-specific generative breakdowns visible only through qualitative analysis, confirming a deep representational gap that prompt engineering cannot resolve. This has significant implications for authentic autistic communication representation. Future research could explore improved alignment strategies to reduce persona-specific generative breakdowns and develop more inclusive models.
Deep Analysis
Background
In recent years, large language models (LLMs) have made significant advances in natural language processing. However, these models may exhibit algorithmic fragility and persona bias when generating autistic communication. Previous studies have focused on model performance in generating neurotypical communication, neglecting the uniqueness of autistic communication.
Core Problem
The core problem is the persona-specific breakdown in LLM-generated autistic communication. Current alignment training may cause models to produce outputs that deviate from source texts under the autistic persona, failing to authentically represent autistic communication.
Innovation
The core innovation is the introduction of a dual-persona rewrite paradigm, revealing persona-specific breakdowns in LLM-generated autistic communication for the first time. A multi-agent qualitative analysis framework systematically analyzes model performance under different personas.
Methodology
- �� Use a dual-persona rewrite paradigm, prompting models to rewrite discourse under autistic and neurotypical personas.
- �� Introduce a multi-agent qualitative analysis framework, analyzing 14,840 rewrite pairs.
- �� Compare model outputs with autistic human annotators' labels, revealing persona-specific generative breakdowns.
Experiments
The experimental design involves using ten LLMs to rewrite autistic discourse under autistic and neurotypical personas. A multi-agent qualitative analysis framework analyzes 14,840 rewrite pairs, revealing persona-specific generative breakdowns.
Results
Results show that autistic persona rewrites diverge significantly in lexical form and affective register compared to neurotypical rewrites, despite equivalent semantic similarity. Most models produce nearly identical outputs across personas, indicating persona-specific generative breakdown.
Applications
The findings can be used to improve LLM performance in generating autistic communication, developing more inclusive models that authentically represent autistic discourse.
Limitations & Outlook
The study's limitations include outputs generated under the autistic persona deviating in lexical and affective aspects, indicating persona-specific bias. Future research could explore improved alignment strategies to reduce persona-specific generative breakdowns.
Plain Language Accessible to non-experts
Imagine a factory where workers process raw materials into products. LLMs are like machines in this factory, converting input text into output. During this process, machines may fail to correctly process certain special materials, like autistic communication. The study finds that when machines handle autistic communication, they often produce biased outputs, not matching the raw materials. It's like factory machines malfunctioning when processing special materials, leading to defective products. To improve this, researchers suggest adjusting the machines to better handle these special materials.
ELI14 Explained like you're 14
Imagine you're playing a role-playing game where you can choose different characters, like a warrior or a wizard. Each character has different skills and missions. In this study, scientists let computer programs play different roles, like autistic and typical roles. They found that when programs play the autistic role, they often mess up, producing content that doesn't quite fit, like a game character using the wrong skills. Scientists want to figure out what's going wrong and improve the programs so they perform better when playing the autistic role.
Glossary
Large Language Model (LLM)
A type of AI model trained on large datasets to generate and understand natural language.
The study used ten LLMs to rewrite autistic discourse.
Persona Bias
The bias exhibited by models when generating text under different personas, leading to inconsistent outputs.
The study reveals persona-specific bias in LLM-generated autistic communication.
Algorithmic Fragility
The tendency of models to err or produce biased outputs when handling specific tasks.
The study finds algorithmic fragility in LLM-generated autistic communication.
Multi-Agent Qualitative Analysis Framework
An analytical method using multiple models to independently analyze data, revealing underlying patterns and biases.
The study used this framework to analyze 14,840 rewrite pairs.
Autistic Communication
The communication style between autistic individuals, often characterized by unique lexical and affective features.
The study prompted models to rewrite autistic communication texts.
Open Questions Unanswered questions from this research
- 1 How to improve alignment strategies to reduce persona-specific generative breakdowns remains an open question.
- 2 The issue of biased outputs under the autistic persona in current models is not fully resolved.
Applications
Immediate Applications
Autistic Communication Generation
Improve LLM performance in generating autistic communication, developing more inclusive models.
Long-term Vision
Inclusive AI
Develop AI capable of authentically representing diverse communication styles, promoting social inclusion.
Abstract
Safety alignment reduces explicitly harmful outputs but inadvertently encodes a sanitized, neuronormative representation of marginalized communication. We investigate this encoding using a dual-persona rewrite paradigm, prompting ten large language models (LLMs) to rewrite naturally occurring autistic discourse from either an autistic or neurotypical persona. We uncover autistic-persona rewrites diverge significantly more in lexical form and affective register than neurotypical rewrites, despite equivalent semantic similarity. Furthermore, most models collapse cross-persona generations into near-identical outputs. To uncover the mechanisms behind this generative breakdown, we introduce a multi-agent qualitative analysis framework. Our results reveal systemic output erasure, stereotyped hallucination, and task-evasive meta-commentary are pervasive failure modes for this task that cluster by alignment strategy rather than parameter scale. Finally, our targeted comparison with autistic human annotators demonstrates that community-insider knowledge produces systematic label reversals relative to LLM classifications. Our findings indicate that current alignment training causes persona-specific generative breakdown visible only through qualitative analysis, confirming a deep representational gap that prompt engineering cannot resolve.