Creating Multilingual Mental Health Dialogue Datasets: Limits of Persona-Based Localization via Nationality and Language
Generated multilingual mental health dialogue datasets by modifying nationality and language parameters, revealing clinical inconsistency issues.
Key Findings
Methodology
The study generated multilingual mental health dialogue datasets by modifying clinical persona nationality and language parameters. Different large language models evaluated depression severity in these datasets, comparing performance against English baselines.
Key Results
- Results showed English models performed best, with accuracy up to 98.21%, while Bengali accuracy dropped to 63.33%.
- Different models showed inconsistent performance in assessing depression severity in non-English texts, especially smaller models.
- Simple replacement of language and nationality failed to maintain clinical signal consistency.
Significance
The study highlights systemic limitations of English-centric personas in multilingual contexts, emphasizing the need for culturally responsive data generation to ensure equitable global mental health systems.
Technical Contribution
The study demonstrates the limitations of persona-based multilingual data generation methods and calls for language-specific evaluations to maintain cross-lingual consistency.
Novelty
First systematic evaluation of clinical consistency in multilingual mental health datasets generated by simple nationality and language parameter replacement.
Limitations
- Simple nationality and language replacement failed to maintain clinical signal consistency, leading to cross-lingual performance disparities.
- Smaller models showed poor performance in non-English contexts, with significant accuracy fluctuations.
Future Work
Future research should focus on culturally responsive data generation, developing more complex persona models to enhance cross-lingual clinical consistency.
AI Executive Summary
Global mental health challenges require high-quality datasets to train and evaluate AI systems. However, existing datasets are primarily English-centric, leading to underrepresentation in other languages. This paper generated multilingual mental health dialogue datasets by modifying persona nationality and language parameters, evaluating different large language models' performance on these datasets. Results showed simple parameter replacement failed to maintain clinical signal consistency, especially in smaller models, revealing significant cross-lingual disparities. The study emphasizes the need for culturally responsive data generation to ensure equitable global mental health systems and calls for language-specific evaluations to maintain cross-lingual consistency.
Deep Analysis
Background
Global mental health issues are increasingly severe, especially in regions with limited medical resources. AI and large language models are seen as potential solutions, but high-quality training and evaluation datasets remain scarce, particularly in non-English contexts.
Core Problem
Existing mental health datasets are primarily English-centric, leading to underrepresentation in other languages, affecting AI systems' performance in multilingual contexts.
Innovation
Generated multilingual mental health dialogue datasets by modifying persona nationality and language parameters, evaluating different large language models' performance on these datasets.
Methodology
- �� Modified persona nationality and language parameters to generate multilingual datasets
- �� Used different large language models to evaluate depression severity in generated datasets
- �� Compared performance against English baselines
Experiments
Experiments used different large language models, including ChatGPT-4o-mini and DeepSeek-V3.2, to evaluate depression severity in generated multilingual datasets. Results showed English models performed best, while other languages showed inconsistent performance.
Results
Results showed English models performed best, with accuracy up to 98.21%, while Bengali accuracy dropped to 63.33%. Different models showed inconsistent performance in assessing depression severity in non-English texts, especially smaller models.
Applications
Findings can be used to develop more complex persona models to enhance cross-lingual clinical consistency, promoting equitable global mental health systems.
Limitations & Outlook
Simple nationality and language replacement failed to maintain clinical signal consistency, leading to cross-lingual performance disparities. Smaller models showed poor performance in non-English contexts, with significant accuracy fluctuations.
Plain Language Accessible to non-experts
Imagine traveling to different countries, trying to communicate in the local language. While you know some basic phrases, you find it difficult to accurately express complex emotions or symptoms. This is similar to the problem found in the study: simple language and nationality replacement fails to maintain clinical signal consistency. Different languages' cultural backgrounds and expressions affect AI models' performance, leading to errors in mental health assessment. Just like in travel, you need a deep understanding of local culture and language to communicate effectively, and AI systems need more complex persona models to improve cross-lingual clinical consistency.
ELI14 Explained like you're 14
Imagine playing a game where your character can travel to different countries. Each country has its own language and culture, and you need to communicate with your character using the local language. But sometimes simple translations can't accurately express the character's emotions or symptoms. That's the problem found in the study: simple language and nationality replacement fails to maintain clinical signal consistency. Different languages' cultural backgrounds and expressions affect AI models' performance, leading to errors in mental health assessment. Just like in the game, you need a deep understanding of local culture and language to communicate effectively, and AI systems need more complex persona models to improve cross-lingual clinical consistency.
Glossary
Large Language Model
An AI model capable of processing and generating natural language text.
Used to evaluate depression severity in generated multilingual mental health datasets.
Persona
A fictional, data-informed character used to simulate user behavior.
Used to generate multilingual mental health dialogue datasets.
Clinical Consistency
Refers to maintaining the same clinical signals across different languages and cultural backgrounds.
Evaluating the clinical consistency of generated multilingual datasets.
Depression Severity
Measures the severity of an individual's depression symptoms.
Using large language models to evaluate depression severity in generated datasets.
Culturally Responsive Data
Data that accurately reflects different cultural backgrounds.
Emphasizing the need for culturally responsive data to ensure equitable global mental health systems.
Open Questions Unanswered questions from this research
- 1 Simple language and nationality replacement failed to maintain clinical signal consistency, future work needs more complex persona models.
- 2 Different languages' cultural backgrounds and expressions affect AI models' performance, future work needs language-specific evaluations.
Applications
Immediate Applications
Multilingual Mental Health Assessment
Develop more complex persona models to improve cross-lingual clinical consistency.
Long-term Vision
Equitable Global Mental Health Systems
Generate culturally responsive data to ensure equitable global mental health systems.
Abstract
AI and large language models (LLMs) have emerged as promising tools to address global mental health challenges. Despite the global nature of these challenges, there remains a critical shortage of high-quality datasets for training and evaluating such systems. To mitigate this gap, researchers increasingly generate synthetic clinical personas to simulate user data and test digital mental health support systems. However, most validated personas rely on English-centric contexts. This paper investigates whether similar persona-based methods can be used to generate multilingual mental health datasets. We modified nationality and language parameters in personas to generate clinical dialogues in Mandarin, Bengali, and Hindi. We then examined how different LLMs perform when evaluating the depression severity of these generated multilingual datasets against the baseline in English. Our findings indicate that just adding nationality and language parameters in personas might not be adequate, as it can introduce clinical inconsistency across languages. LLM judge models often exhibit inaccuracies in assessing depression severity in non-English texts, with performance varying across different models. This exposes the systemic limitations of applying English-centric personas to multilingual contexts. Ultimately, our work highlights the urgent need for culturally responsive data generation to ensure equitable mental health systems globally.