Do Large Language Models Truly Understand Cross-cultural Differences?

TL;DR

SAGE benchmark evaluates LLMs' cross-cultural understanding via 9 dimensions and 4,530 tasks, exposing systematic deficiencies.

cs.CL 🔴 Advanced 2025-12-08 36 views
Shiwei Guo Sihang Jiang Qianxi He Yanghua Xiao Jiaqing Liang Bi Yude Minggui He Shimin Tao Li Zhang
LLMs cross-cultural understanding SAGE benchmark deep cultural reasoning multilingual evaluation

Key Findings

Methodology

SAGE benchmark aligns Cross-Cultural Core Concepts (CCCs) and generative task design, covering 9 dimensions, 4 scenario types, and 15 real-world contexts with 4,530 tasks. It supports multilingual expansion.

Key Results

  • Result 1: GPT-4o achieved the highest accuracy (41%) in Chinese tasks but only 40% in Spanish, revealing language imbalance.
  • Result 2: Small models like LLaMA2.5-7B scored as low as 2% accuracy in abstract dimensions (e.g., metaphysics), highlighting deficiencies in deep cultural reasoning.
  • Result 3: SAGE reveals systematic failures in cross-cultural reasoning, particularly in abstract domains like values and cognition.

Significance

SAGE fills a critical gap in evaluating LLMs' cross-cultural understanding, providing a comprehensive framework for assessing real-world multilingual applications. This is crucial for improving AI's global usability.

Technical Contribution

Introduced the first benchmark using CCCs for dynamic scenario-based evaluation, avoiding monocultural assumptions and supporting multilingual scalability. It highlights deficiencies in deep cultural reasoning.

Novelty

SAGE is the first benchmark to evaluate cross-cultural understanding through dynamic scenarios and CCC alignment, surpassing traditional benchmarks focused on cultural trivia.

Limitations

  • Limitation 1: Currently supports only Chinese and Spanish; further language expansion needs validation.
  • Limitation 2: Deep cultural reasoning evaluation relies on manually designed questions and scoring.
  • Limitation 3: Performance in abstract dimensions may be biased by training data's cultural skew.

Future Work

Future work includes expanding to more languages, automating deep cultural reasoning evaluation, and improving LLMs' cross-cultural capabilities through training.

AI Executive Summary

Large Language Models (LLMs) have shown strong multilingual performance, but their ability to understand cross-cultural differences remains underexplored. Existing benchmarks lack contextual scenarios, cross-cultural concept mapping, and deep cultural reasoning. To address this, researchers introduced the SAGE benchmark, which evaluates LLMs using Cross-Cultural Core Concepts (CCCs) aligned with generative task design. SAGE spans 9 dimensions, 4 scenario types, and 15 real-world contexts, comprising 4,530 tasks.

Experiments reveal that while GPT-4o performs best in Chinese tasks (41% accuracy), it slightly underperforms in Spanish (40%), exposing language imbalances. Smaller models like LLaMA2.5-7B struggle in abstract dimensions, with accuracies as low as 2%, highlighting significant gaps in deep cultural reasoning.

SAGE not only fills a critical gap in cross-cultural evaluation but also provides a roadmap for improving LLMs in global applications. However, its current scope is limited to Chinese and Spanish, necessitating future expansion to other languages and automated evaluation methods for deeper cultural reasoning.

Deep Analysis

Background

Recent advancements in LLMs have significantly improved multilingual tasks. However, cross-cultural understanding—a critical competency for real-world applications—remains underexplored. Existing benchmarks primarily test cultural trivia, lacking contextual and reasoning depth.

Core Problem

The core issue lies in LLMs' systematic deficiencies in cross-cultural reasoning, particularly in abstract dimensions like values and cognition. Understanding cultural differences requires more than multilingual capabilities; it demands deep reasoning about cultural contexts.

Innovation

SAGE introduces a novel evaluation framework based on CCC alignment and generative task design. Key innovations include:

  • �� A 9-dimension framework grounded in cultural theory.
  • �� 15 real-world scenarios with 4,530 tasks.
  • �� Multilingual scalability, initially supporting Chinese and Spanish.

Methodology

SAGE's construction involves:

  • �� Defining CCCs across 9 categories, including values and cognition.
  • �� Designing 4 scenario types (e.g., conflict resolution, communication).
  • �� Creating multiple-choice, true/false, and short-answer tasks.
  • �� Validating tasks through expert review and iterative refinement.

Experiments

Experiments tested models like LLaMA2.5-7B, Qwen3-14B, and GPT-4o in Chinese and Spanish. Metrics included accuracy, cross-cultural stability, and dimension-specific performance. Ablation studies analyzed contributions of different components.

Results

GPT-4o achieved 41% accuracy in Chinese tasks but only 40% in Spanish, revealing language imbalances. Small models scored as low as 2% in abstract dimensions, exposing deficiencies in deep cultural reasoning. Language disparities were consistent across models.

Applications

SAGE can evaluate and improve LLMs for global applications, including cross-cultural communication, international business, and multilingual education.

Limitations & Outlook

Currently limited to Chinese and Spanish, with manual question design for deep reasoning. Future work should expand language coverage and automate evaluation methods.

Plain Language Accessible to non-experts

Imagine attending an international food festival where each culture has unique dining rules. An AI assistant needs to know that in China, chopsticks shouldn't be stuck upright in rice, while in Spain, both hands should remain visible on the table. SAGE tests whether AI can not only recall these rules but also understand their cultural significance.

ELI14 Explained like you're 14

Think of AI as a new friend at a global culture fair. It needs to learn that in China, you can't tap chopsticks on a bowl, and in Spain, you should keep your hands on the table. SAGE is like a test to see if this friend can follow the rules and understand why they matter!

Glossary

SAGE Benchmark

A benchmark for evaluating LLMs' cross-cultural understanding using CCC alignment and generative tasks.

Used to test models' reasoning in real-world cultural contexts.

Cross-Cultural Core Concepts (CCCs)

Key concepts significant across cultures, defined by cultural studies.

Core building blocks of the SAGE benchmark.

Deep Cultural Reasoning

The ability to understand values, cognition, and norms within cultural contexts.

Evaluated as a critical dimension in SAGE.

Cultural Vacancy

A concept in one culture with no direct equivalent in another.

Preserved in SAGE to test adaptability.

Multilingual Scalability

The ability to extend benchmarks to other languages with minimal manual effort.

SAGE initially supports Chinese and Spanish.

Open Questions Unanswered questions from this research

  • 1 How can SAGE be expanded to more languages while ensuring cultural fairness?
  • 2 What automated methods can improve deep cultural reasoning evaluation?

Applications

Immediate Applications

Cross-Cultural Communication

Helps businesses and individuals navigate cultural differences in international interactions.

Multilingual Education

Provides cultural context for language learners, enhancing learning outcomes.

Long-term Vision

Global AI Systems

Develop AI systems adaptable to diverse cultural environments, advancing global commerce and society.

Abstract

In recent years, large language models (LLMs) have demonstrated strong performance on multilingual tasks. Given its wide range of applications, cross-cultural understanding capability is a crucial competency. However, existing benchmarks for evaluating whether LLMs genuinely possess this capability suffer from three key limitations: a lack of contextual scenarios, insufficient cross-cultural concept mapping, and limited deep cultural reasoning capabilities. To address these gaps, we propose SAGE, a scenario-based benchmark built via cross-cultural core concept alignment and generative task design, to evaluate LLMs' cross-cultural understanding and reasoning. Grounded in cultural theory, we categorize cross-cultural capabilities into nine dimensions. Using this framework, we curated 210 core concepts and constructed 4530 test items across 15 specific real-world scenarios, organized under four broader categories of cross-cultural situations, following established item design principles. The SAGE dataset supports continuous expansion, and experiments confirm its transferability to other languages. It reveals model weaknesses across both dimensions and scenarios, exposing systematic limitations in cross-cultural reasoning. While progress has been made, LLMs are still some distance away from reaching a truly nuanced cross-cultural understanding. In compliance with the anonymity policy, we include data and code in the supplement materials. In future versions, we will make them publicly available online.

cs.CL