SurakshaEval: An Indic Safety Benchmark for Multilingual LLMs

TL;DR

SurakshaEval evaluates multilingual LLM safety across 10 Indic languages, revealing cultural sensitivity issues.

cs.CL 🔴 Advanced 2026-08-08 3 views
Debopriyo Banerjee Kapil Rajesh Kavitha Angana Borah Xudong Han Yuxia Wang Parameswari Krishnamurthy Utkarsh Agarwal Atharva Kulkarni Swaran Lata Ayush Munot Dhruv Sahnan Aaryamonvikram Singh Preslav Nakov Monojit Choudhury
multilingual safety evaluation Indic languages cultural sensitivity large language models

Key Findings

Methodology

SurakshaEval is a novel safety benchmark covering 10 major Indic languages and English. It uses human-written scenario prompts to evaluate multilingual LLMs on region and language-specific safety risks. PoLL evaluators are used for data verification and filtering to ensure prompt accuracy and cultural fidelity.

Key Results

  • Result 1: Even strong multilingual LLMs perform poorly in nuanced safety requirements for Indic languages, especially in native scripts.
  • Result 2: Common failure modes include over-refusal, missed implicit bias detection, and insufficient contextual awareness in regionally sensitive settings.
  • Result 3: Evaluation of 27 LLMs revealed difficulties in maintaining consistent performance across all safety categories.

Significance

This research fills a gap in existing safety evaluation datasets for multilingual and cultural contexts, particularly for India's diversity. By highlighting LLM safety challenges in Indic languages, it advances the development of AI systems that operate securely, ethically, and in alignment with diverse societal values.

Technical Contribution

SurakshaEval provides a systematic safety evaluation framework incorporating region-specific data and structured assessment protocols. By introducing PoLL evaluators, it enhances prompt verification and filtering, ensuring data accuracy and cultural applicability.

Novelty

SurakshaEval is the first safety benchmark specifically targeting Indic languages and cultural contexts, offering comprehensive evaluation of multilingual LLMs in regional sensitivity.

Limitations

  • Limitation 1: Annotation bias may exist due to dataset complexity and diversity.
  • Limitation 2: Limited to 10 Indic languages, not covering all official languages.

Future Work

Future work will expand to more Indic languages and other cultural contexts, further refining the safety evaluation framework and exploring LLM performance in diverse social environments.

AI Executive Summary

Existing large language model (LLM) safety evaluation datasets predominantly focus on English and Western contexts, overlooking linguistic diversity and culturally grounded safety risks in other languages. SurakshaEval addresses this gap by introducing a novel safety benchmark explicitly designed for evaluating multilingual LLMs in 10 major Indic languages and English. The benchmark includes both generic prompts common across India and region- and language-specific prompts that capture localized sociocultural sensitivities.

Benchmarking a broad range of state-of-the-art LLMs on SurakshaEval reveals that even strong multilingual LLMs struggle with nuanced safety requirements in Indic languages, particularly in native scripts. Common failure modes include over-refusal, missed detection of implicit bias, and insufficient contextual awareness in regionally sensitive settings.

These findings highlight the urgent need for safety evaluation frameworks that incorporate region-specific data and structured assessment protocols, enabling the development and deployment of AI systems that operate securely, ethically, and in alignment with diverse societal values. SurakshaEval provides a systematic framework for this purpose, enhancing prompt verification and filtering through the introduction of PoLL evaluators, ensuring data accuracy and cultural fidelity.

Deep Analysis

Background

As large language models (LLMs) expand into multilingual settings, concerns about the safety and reliability of their generated content grow. Existing safety evaluation datasets focus primarily on English and Western contexts, neglecting the linguistic diversity and cultural safety risks in other languages. India's linguistic diversity and cultural complexity pose challenges for establishing robust safety benchmarks in these environments.

Core Problem

There is a significant gap in existing safety evaluation datasets for multilingual and cultural contexts, particularly in India's diverse setting. This leads to inadequate assessment of LLM performance in these environments, failing to effectively capture region and language-specific safety risks.

Innovation

SurakshaEval introduces a novel safety benchmark specifically targeting India's multilingual and cultural contexts. It uses human-written scenario prompts to evaluate multilingual LLMs on region and language-specific safety risks. PoLL evaluators are used for data verification and filtering to ensure prompt accuracy and cultural fidelity.

Methodology

  • �� Design a safety benchmark covering 10 Indic languages and English.
  • �� Use human-written scenario prompts to cover region and language-specific safety risks.
  • �� Employ PoLL evaluators for data verification and filtering, ensuring prompt accuracy and cultural fidelity.
  • �� Benchmark 27 state-of-the-art LLMs to assess their performance across different safety categories.

Experiments

The experimental design includes evaluating 27 LLMs across 10 Indic languages and English. Using the SurakshaEval benchmark, the models are assessed on region and language-specific safety risks. PoLL evaluators are used for data verification and filtering to ensure prompt accuracy and cultural fidelity.

Results

The results show that even strong multilingual LLMs perform poorly in nuanced safety requirements for Indic languages, especially in native scripts. Common failure modes include over-refusal, missed detection of implicit bias, and insufficient contextual awareness in regionally sensitive settings.

Applications

SurakshaEval can be used to evaluate and improve the safety of multilingual LLMs in Indic language and cultural contexts, aiding in the development of safer and more ethical AI systems.

Limitations & Outlook

Annotation bias may exist due to dataset complexity and diversity. Additionally, SurakshaEval is limited to 10 Indic languages, not covering all official languages. Future work will expand to more Indic languages and other cultural contexts.

Plain Language Accessible to non-experts

Imagine you're in a multilingual Indian market, with each stall representing a language. SurakshaEval acts like a market manager, ensuring each stall is safe and free from offensive content. The manager uses a set of standards to evaluate whether each stall's content meets the market's cultural and linguistic standards. This process is similar to how SurakshaEval evaluates multilingual LLMs across different Indic languages, ensuring they consider cultural sensitivity and safety when generating content.

ELI14 Explained like you're 14

Imagine you're playing a multilingual Indian game, with each level representing a language. SurakshaEval is like the game's guardian, ensuring each level is free from dangerous content. The guardian uses a set of rules to check if each level's content meets the game's cultural and linguistic standards. This is like how SurakshaEval evaluates multilingual LLMs across different Indic languages, ensuring they consider cultural sensitivity and safety when generating content.

Glossary

Large Language Model (LLM)

A large machine learning model capable of understanding and generating natural language.

Used in the paper to evaluate its safety in multilingual and cultural contexts.

Safety Benchmark

A standard or framework used to evaluate a model's safety in a specific environment.

SurakshaEval serves as a new safety benchmark for evaluating multilingual LLMs.

PoLL Evaluator

An evaluation tool used for data verification and filtering, ensuring prompt accuracy and cultural fidelity.

Used in SurakshaEval for data verification and filtering.

Region-Specific Risk

Safety risks related to the cultural and linguistic context of a specific region.

Used in SurakshaEval to evaluate multilingual LLM performance.

Cultural Sensitivity

Understanding and respecting specific cultural contexts to avoid offensive content.

Used in SurakshaEval to evaluate the safety of multilingual LLMs.

Open Questions Unanswered questions from this research

  • 1 How can SurakshaEval be expanded to more Indic languages and other cultural contexts?
  • 2 How to reduce annotation bias in the dataset to ensure accurate evaluation?

Applications

Immediate Applications

Multilingual LLM Safety Evaluation

SurakshaEval can be used to evaluate and improve the safety of multilingual LLMs in Indic language and cultural contexts.

Long-term Vision

Global Multilingual Safety Framework

SurakshaEval's framework can be extended to other multilingual and cultural contexts, aiding in the development of safer global AI systems.

Abstract

Existing safety evaluation datasets for large language models (LLMs) predominantly focus on English and Western contexts, often overlooking the linguistic diversity and culturally grounded safety risks present in other languages. To address this gap, we introduce SurakshaEval, a novel safety benchmark composed of human-written prompts spanning real-world scenarios, explicitly designed for ten major Indian languages - Assamese, Bengali, Gujarati, Hindi, Kannada, Malayalam, Marathi, Punjabi, Tamil, and Telugu, along with English. SurakshaEval includes both generic prompts common across India and region- and language-specific prompts that capture localized sociocultural sensitivities. We benchmark a broad range of state-of-the-art LLMs on SurakshaEval, establish baseline safety performance, and identify recurring failure modes, including over-refusal, missed detection of implicit bias, and insufficient contextual awareness in regionally sensitive settings. Our results show that even strong multilingual LLMs struggle to reliably meet nuanced safety requirements when operating in Indic languages, particularly in native scripts. These findings highlight the urgent need for safety evaluation frameworks that incorporate region-specific data and structured assessment protocols, enabling the development and deployment of AI systems that operate securely, ethically, and in alignment with diverse societal values. Our code and data are available at https://github.com/debobanerjee/SurakshaEval. Warning: This paper contains text that may be offensive or unsafe.

cs.CL