Creating Multilingual Mental Health Dialogue Datasets: Limits of Persona-Based Localization via Nationality and Language

TL;DR

Generated multilingual mental health dialogue datasets by modifying nationality and language parameters, revealing clinical inconsistency issues.

cs.CL 🔴 Advanced 2026-06-18 1 views
Yunkai Xu Saeed Abdullah
mental health multilingual large language models datasets clinical consistency

Key Findings

Methodology

The study generated multilingual mental health dialogue datasets by modifying clinical persona nationality and language parameters. Different large language models evaluated depression severity in these datasets, comparing performance against English baselines.

Key Results

  • Results showed English models performed best, with accuracy up to 98.21%, while Bengali accuracy dropped to 63.33%.
  • Different models showed inconsistent performance in assessing depression severity in non-English texts, especially smaller models.
  • Simple replacement of language and nationality failed to maintain clinical signal consistency.

Significance

The study highlights systemic limitations of English-centric personas in multilingual contexts, emphasizing the need for culturally responsive data generation to ensure equitable global mental health systems.

Technical Contribution

The study demonstrates the limitations of persona-based multilingual data generation methods and calls for language-specific evaluations to maintain cross-lingual consistency.

Novelty

First systematic evaluation of clinical consistency in multilingual mental health datasets generated by simple nationality and language parameter replacement.

Limitations

  • Simple nationality and language replacement failed to maintain clinical signal consistency, leading to cross-lingual performance disparities.
  • Smaller models showed poor performance in non-English contexts, with significant accuracy fluctuations.

Future Work

Future research should focus on culturally responsive data generation, developing more complex persona models to enhance cross-lingual clinical consistency.

AI Executive Summary

Global mental health challenges require high-quality datasets to train and evaluate AI systems. However, existing datasets are primarily English-centric, leading to underrepresentation in other languages. This paper generated multilingual mental health dialogue datasets by modifying persona nationality and language parameters, evaluating different large language models' performance on these datasets. Results showed simple parameter replacement failed to maintain clinical signal consistency, especially in smaller models, revealing significant cross-lingual disparities. The study emphasizes the need for culturally responsive data generation to ensure equitable global mental health systems and calls for language-specific evaluations to maintain cross-lingual consistency.

Deep Analysis

Background

Global mental health issues are increasingly severe, especially in regions with limited medical resources. AI and large language models are seen as potential solutions, but high-quality training and evaluation datasets remain scarce, particularly in non-English contexts.

Core Problem

Existing mental health datasets are primarily English-centric, leading to underrepresentation in other languages, affecting AI systems' performance in multilingual contexts.

Innovation

Generated multilingual mental health dialogue datasets by modifying persona nationality and language parameters, evaluating different large language models' performance on these datasets.

Methodology

  • �� Modified persona nationality and language parameters to generate multilingual datasets
  • �� Used different large language models to evaluate depression severity in generated datasets
  • �� Compared performance against English baselines

Experiments

Experiments used different large language models, including ChatGPT-4o-mini and DeepSeek-V3.2, to evaluate depression severity in generated multilingual datasets. Results showed English models performed best, while other languages showed inconsistent performance.

Results

Results showed English models performed best, with accuracy up to 98.21%, while Bengali accuracy dropped to 63.33%. Different models showed inconsistent performance in assessing depression severity in non-English texts, especially smaller models.

Applications

Findings can be used to develop more complex persona models to enhance cross-lingual clinical consistency, promoting equitable global mental health systems.

Limitations & Outlook

Simple nationality and language replacement failed to maintain clinical signal consistency, leading to cross-lingual performance disparities. Smaller models showed poor performance in non-English contexts, with significant accuracy fluctuations.

Plain Language Accessible to non-experts

Imagine traveling to different countries, trying to communicate in the local language. While you know some basic phrases, you find it difficult to accurately express complex emotions or symptoms. This is similar to the problem found in the study: simple language and nationality replacement fails to maintain clinical signal consistency. Different languages' cultural backgrounds and expressions affect AI models' performance, leading to errors in mental health assessment. Just like in travel, you need a deep understanding of local culture and language to communicate effectively, and AI systems need more complex persona models to improve cross-lingual clinical consistency.

ELI14 Explained like you're 14

Imagine playing a game where your character can travel to different countries. Each country has its own language and culture, and you need to communicate with your character using the local language. But sometimes simple translations can't accurately express the character's emotions or symptoms. That's the problem found in the study: simple language and nationality replacement fails to maintain clinical signal consistency. Different languages' cultural backgrounds and expressions affect AI models' performance, leading to errors in mental health assessment. Just like in the game, you need a deep understanding of local culture and language to communicate effectively, and AI systems need more complex persona models to improve cross-lingual clinical consistency.

Glossary

Large Language Model

An AI model capable of processing and generating natural language text.

Used to evaluate depression severity in generated multilingual mental health datasets.

Persona

A fictional, data-informed character used to simulate user behavior.

Used to generate multilingual mental health dialogue datasets.

Clinical Consistency

Refers to maintaining the same clinical signals across different languages and cultural backgrounds.

Evaluating the clinical consistency of generated multilingual datasets.

Depression Severity

Measures the severity of an individual's depression symptoms.

Using large language models to evaluate depression severity in generated datasets.

Culturally Responsive Data

Data that accurately reflects different cultural backgrounds.

Emphasizing the need for culturally responsive data to ensure equitable global mental health systems.

Open Questions Unanswered questions from this research

  • 1 Simple language and nationality replacement failed to maintain clinical signal consistency, future work needs more complex persona models.
  • 2 Different languages' cultural backgrounds and expressions affect AI models' performance, future work needs language-specific evaluations.

Applications

Immediate Applications

Multilingual Mental Health Assessment

Develop more complex persona models to improve cross-lingual clinical consistency.

Long-term Vision

Equitable Global Mental Health Systems

Generate culturally responsive data to ensure equitable global mental health systems.

Abstract

AI and large language models (LLMs) have emerged as promising tools to address global mental health challenges. Despite the global nature of these challenges, there remains a critical shortage of high-quality datasets for training and evaluating such systems. To mitigate this gap, researchers increasingly generate synthetic clinical personas to simulate user data and test digital mental health support systems. However, most validated personas rely on English-centric contexts. This paper investigates whether similar persona-based methods can be used to generate multilingual mental health datasets. We modified nationality and language parameters in personas to generate clinical dialogues in Mandarin, Bengali, and Hindi. We then examined how different LLMs perform when evaluating the depression severity of these generated multilingual datasets against the baseline in English. Our findings indicate that just adding nationality and language parameters in personas might not be adequate, as it can introduce clinical inconsistency across languages. LLM judge models often exhibit inaccuracies in assessing depression severity in non-English texts, with performance varying across different models. This exposes the systemic limitations of applying English-centric personas to multilingual contexts. Ultimately, our work highlights the urgent need for culturally responsive data generation to ensure equitable mental health systems globally.

cs.CL cs.AI cs.HC