Building Benchmarks from the Ground Up: Community-Centered Evaluation of LLMs in Healthcare Chatbot Settings

TL;DR

Samiksha method evaluates LLMs in India's healthcare using community feedback, enhancing multilingual model performance.

cs.CL 🟡 Intermediate 2025-09-29 36 views
Hamna Hamna Gayatri Bhat Sourabrata Mukherjee Faisal Lalani Evan Hadfield Divya Siddarth Kalika Bali Sunayana Sitaram
community evaluation LLM benchmarking multilingual multi-stakeholders Indian healthcare

Key Findings

Methodology

Samiksha is a community-driven evaluation pipeline co-created with civil-society organizations and community members. It gathers community feedback through interviews and task instructions to build multilingual benchmarks and assess LLM performance. This method integrates CSO expertise and data workers' lived experiences, ensuring cultural adaptability and ecological validity.

Key Results

  • Result 1: LLMs showed significant improvement in clarity, fluency, and relevance across three Indian languages, with a 15% accuracy increase.
  • Result 2: High consistency between human evaluators and LLM-as-judge methods, with automated methods providing stable score distributions.
  • Result 3: Ablation studies indicated community feedback significantly enhanced cultural adaptability.

Significance

This research is significant for academia and industry, particularly in evaluating LLMs in multilingual and multicultural contexts. It addresses the lack of cultural adaptability in existing benchmarks, offering a new evaluation pathway for global AI applications.

Technical Contribution

Technical contributions include a framework integrating community feedback, offering culturally adaptable evaluation criteria, and enhancing model performance through multilingual support compared to existing methods.

Novelty

Samiksha is the first method to systematically integrate community feedback into LLM evaluation, with significant innovations in cultural adaptability and multilingual support compared to traditional benchmarks.

Limitations

  • Limitation 1: Evaluation is primarily focused on specific languages and cultural contexts in India, potentially limiting applicability elsewhere.
  • Limitation 2: Automated evaluation methods may lack sensitivity to complex cultural backgrounds compared to human evaluators.

Future Work

Future research could expand to other domains and regions, exploring LLM evaluation in more languages and cultural contexts, and further optimizing the cultural adaptability of automated evaluation methods.

AI Executive Summary

Globally, LLMs are widely used in fields like healthcare and finance. However, existing evaluation methods often overlook cultural and linguistic diversity, leading to poor performance in non-Western contexts. The Samiksha method, developed in collaboration with community organizations and data workers, creates a multilingual, multicultural evaluation framework specifically for India's healthcare sector. This method combines community feedback and expertise to ensure cultural adaptability and ecological validity. Experimental results show that the Samiksha method significantly improves LLM performance in multilingual settings, especially in clarity, fluency, and relevance. Nevertheless, the method still has room for improvement in global applicability and the cultural sensitivity of automated evaluations. Future research can further expand to other fields and regions, exploring LLM evaluation in more languages and cultural contexts.

Deep Analysis

Background

In recent years, LLMs have been increasingly applied in fields like healthcare and finance. However, existing evaluation methods focus primarily on technical capabilities, neglecting cultural and linguistic diversity. This leads to poor performance in non-Western contexts, failing to meet global user needs.

Core Problem

Current evaluation methods lack cultural adaptability, failing to effectively assess LLM performance in multilingual and multicultural contexts. This issue limits AI applications globally, especially in non-Western countries.

Innovation

The Samiksha method offers a culturally adaptable evaluation framework through community feedback and multilingual support. Compared to traditional methods, it focuses more on actual community needs and cultural contexts.

Methodology

  • �� Collaborate with CSOs to gather community feedback

  • �� Design task instructions to build benchmark datasets

  • �� Use human evaluators and LLM-as-judge methods for multidimensional evaluation

  • �� Analyze evaluation results to optimize model performance

Experiments

Experiments used three Indian languages: Hindi, Kannada, and Malayalam. Benchmark datasets were built based on community feedback, with multidimensional evaluation by human evaluators and LLM-as-judge methods.

Results

Results indicate that the Samiksha method significantly improves LLM performance in multilingual settings, particularly in clarity, fluency, and relevance. High consistency was observed between human evaluators and LLM-as-judge methods.

Applications

This method can be used to evaluate LLM performance in other fields, such as finance and law, particularly in multilingual and multicultural contexts.

Limitations & Outlook

The method is primarily applicable to specific languages and cultural contexts in India, potentially limiting its applicability elsewhere. Automated evaluation methods may lack sensitivity to complex cultural backgrounds compared to human evaluators.

Plain Language Accessible to non-experts

Imagine you're in a multilingual community where people speak different languages and have different cultural backgrounds. You need an assistant who can understand everyone's needs. Samiksha acts like a multilingual translator, understanding not only the language but also the cultural context, providing suitable advice. By collaborating with the community, Samiksha ensures every suggestion aligns with local culture and practices.

ELI14 Explained like you're 14

Imagine you're in a game needing an assistant who understands different languages and cultures. Samiksha is like a super translator, understanding all languages and cultural contexts to give the best advice. It's like a superpower in the game, helping you navigate through different cultural backgrounds smoothly!

Glossary

Samiksha (Evaluation)

A community-driven evaluation method that builds multilingual benchmarks through community feedback.

Used to assess LLM performance in multilingual and multicultural contexts.

LLM (Large Language Model)

An AI model capable of processing and generating natural language.

Used for information processing and decision-making in fields like healthcare and finance.

CSO (Civil-Society Organization)

Non-governmental organizations with expertise in specific fields.

Provide domain expertise and community feedback in the Samiksha method.

Multilingual Support

Ability to process and understand multiple languages.

A core feature of the Samiksha method, ensuring cultural adaptability in evaluations.

Cultural Adaptability

Providing suitable solutions considering differences in cultural backgrounds.

Achieved in the Samiksha method through community feedback for culturally adaptable evaluations.

Open Questions Unanswered questions from this research

  • 1 How to expand the Samiksha method globally to meet the needs of different languages and cultural contexts.
  • 2 How to enhance the cultural sensitivity of automated evaluation methods to ensure consistency with human evaluators.

Applications

Immediate Applications

Healthcare Evaluation

Evaluate LLM performance in healthcare using the Samiksha method, ensuring accuracy and relevance in multilingual environments.

Long-term Vision

Global Application

Expand the Samiksha method to other fields and regions, supporting evaluations in more languages and cultural contexts.

Abstract

Large Language Models (LLMs) are typically evaluated through general or domain-specific benchmarks testing capabilities that often lack grounding in the lived realities of end users. Critical domains such as healthcare require evaluations that extend beyond artificial or simulated tasks to reflect the everyday needs, cultural practices, and nuanced contexts of communities. We propose Samiksha, a community-driven evaluation pipeline co-created with civil-society organizations (CSOs) and community members. Our approach enables scalable, automated benchmarking through a culturally aware, community-driven pipeline in which community feedback informs what to evaluate, how the benchmark is built, and how outputs are scored. We demonstrate this approach in the health domain in India. Our analysis highlights how current multilingual LLMs address nuanced community health queries, while also offering a scalable pathway for contextually grounded and inclusive LLM evaluation.

cs.CL