GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models

TL;DR

GSM-Plus-BN evaluates LLMs on Bengali math reasoning with perturbations; GPT-OSS-20B achieves 96.08% accuracy on seed questions under Standard Prompting.

cs.CL 🔴 Advanced 2026-07-15 33 views
Bidyarthi Paul Nahida Jannat Mayouree Md. Asif Karim Sagar Chandra Nath Swastika Kundu
mathematical reasoning LLMs Bengali perturbation robustness Chain-of-Thought

Key Findings

Methodology

The study introduces GSM-Plus-BN, a Bengali mathematical reasoning dataset derived from GSM-Plus, with 1,000 seed questions and 8,000 perturbations. Six open-source LLMs, including Qwen3-32B and Llama-3.3-70B, were evaluated using Standard Prompting and Chain-of-Thought (CoT) Prompting.

Key Results

  • GPT-OSS-20B achieved 96.08% accuracy on seed questions under Standard Prompting, the highest among tested models.
  • Llama-3.3-70B and GPT-OSS-120B demonstrated superior robustness across multiple perturbation types.
  • CoT prompting significantly improved reasoning for most models, but a performance gap remains compared to English benchmarks.

Significance

This research fills a critical gap in Bengali mathematical reasoning by providing the first systematic evaluation of LLMs in this low-resource language. It highlights the challenges of achieving robustness and true understanding in non-English contexts.

Technical Contribution

Key contributions include the GSM-Plus-BN dataset covering eight perturbation types, systematic evaluation of six LLMs, and insights into the effectiveness of CoT prompting for Bengali reasoning tasks.

Novelty

This is the first perturbation-based Bengali mathematical reasoning benchmark, enabling cross-lingual evaluation and addressing a long-standing gap in low-resource language research.

Limitations

  • Models perform significantly worse in Bengali compared to English, indicating limited adaptation to low-resource languages.
  • The dataset does not include more complex mathematical problems.
  • Experiments were limited to six open-source models, excluding proprietary systems like GPT-4.

Future Work

Future work could expand the dataset to include more complex problems, explore specialized pretraining for low-resource languages, and refine CoT prompting strategies.

AI Executive Summary

Mathematical reasoning is a key benchmark for evaluating the cognitive capabilities of large language models (LLMs). However, most research has focused on high-resource languages like English, leaving low-resource languages like Bengali underexplored. GSM-Plus-BN addresses this gap by introducing a perturbation-based Bengali mathematical reasoning dataset with 1,000 seed questions and 8,000 perturbations.

The study evaluates six open-source LLMs, including Qwen3-32B and Llama-3.3-70B, using Standard Prompting and Chain-of-Thought (CoT) Prompting. Results show that GPT-OSS-20B achieved the highest accuracy of 96.08% on seed questions under Standard Prompting, while Llama-3.3-70B and GPT-OSS-120B exhibited greater robustness to perturbations. CoT prompting significantly improved reasoning, though all models performed worse in Bengali compared to English benchmarks.

By providing GSM-Plus-BN, this research lays the groundwork for future studies in Bengali mathematical reasoning. It underscores the need for improved LLMs tailored to low-resource languages and highlights the potential of CoT prompting to enhance reasoning capabilities. Future work should focus on expanding the dataset and developing more robust models for non-English languages.

Deep Analysis

Background

Mathematical reasoning is a critical area for evaluating LLMs. Benchmarks like GSM8K and MATH have advanced English-language research, but low-resource languages like Bengali remain underrepresented. Bengali, spoken by over 300 million people, lacks robust datasets and evaluation frameworks for complex reasoning tasks.

Core Problem

Existing LLMs lack systematic evaluation for mathematical reasoning in low-resource languages like Bengali, especially under input perturbations. This gap limits the equitable development of AI systems for diverse linguistic communities.

Innovation

GSM-Plus-BN is the first perturbation-based Bengali mathematical reasoning dataset, featuring eight perturbation types such as numerical substitution and problem rephrasing. The study also evaluates six LLMs, providing insights into their robustness and reasoning capabilities in Bengali.

Methodology

  • �� Dataset construction: Derived 1,000 seed questions and 8,000 perturbations from GSM-Plus.
  • �� Model selection: Evaluated six LLMs, including Qwen3-32B and Llama-3.3-70B.
  • �� Prompting strategies: Used Standard Prompting and Chain-of-Thought (CoT) Prompting.
  • �� Evaluation: Analyzed model performance across perturbation types and compared to English benchmarks.

Experiments

The GSM-Plus-BN dataset includes 9,000 samples. Six LLMs were tested under Standard and CoT prompting. Performance was analyzed across eight perturbation types to assess robustness and reasoning capabilities.

Results

GPT-OSS-20B achieved 96.08% accuracy on seed questions under Standard Prompting. Llama-3.3-70B and GPT-OSS-120B showed higher robustness to perturbations. CoT prompting improved reasoning but did not close the gap with English benchmarks.

Applications

GSM-Plus-BN can be used to evaluate LLMs in low-resource languages, supporting educational tools, translation systems, and AI applications in diverse linguistic settings.

Limitations & Outlook

Models performed worse in Bengali compared to English. The dataset lacks complex problems, and experiments were limited to six open-source models.

Plain Language Accessible to non-experts

Imagine testing a student by rephrasing math problems or using a different language. GSM-Plus-BN does this for AI models, checking if they truly understand the problem or just memorize patterns. This helps researchers see how well AI adapts to new challenges.

ELI14 Explained like you're 14

Think of a math test where the teacher changes the numbers or rewrites the questions in another language. GSM-Plus-BN is like that for AI! It checks if the AI can still solve the problem even when things are switched up. Cool, right?

Glossary

GSM-Plus-BN

A Bengali math reasoning dataset with perturbations to test LLM robustness.

Used to evaluate LLM performance in low-resource languages.

Chain-of-Thought Prompting

A strategy encouraging step-by-step reasoning to improve accuracy.

Used to enhance reasoning in Bengali tasks.

Perturbation Types

Modifications to questions, like numerical substitution or rephrasing.

Used to test model robustness to input variations.

Low-Resource Language

Languages with limited NLP resources, like Bengali.

Target language for this study.

Robustness

The ability of models to maintain performance under input variations.

Evaluated through GSM-Plus-BN perturbations.

Open Questions Unanswered questions from this research

  • 1 How can LLMs improve reasoning in low-resource languages like Bengali?
  • 2 What new perturbation types could better test model robustness?

Applications

Immediate Applications

Educational Tools

Develop AI tools for Bengali math education, aiding students in learning.

Multilingual AI

Enhance AI performance in low-resource languages for broader accessibility.

Long-term Vision

Global AI Systems

Create AI systems that work seamlessly across all languages, breaking barriers.

Abstract

The evaluation of mathematical reasoning in large language models (LLMs) has predominantly focused on high-resource languages like English. This has created a significant barrier to the equitable development and deployment of AI in linguistically diverse regions such as Bangladesh, where over 230 million people speak Bengali. Despite this global significance, there has been minimal prior work on mathematical reasoning in Bengali and no existing research that systematically benchmarks a perturbated Bengali mathematical dataset, leaving a critical void in assessing model robustness and true comprehension beyond pattern recognition. This study addresses this gap by introducing GSM-Plus-BN, a novel perturbated Bengali mathematical dataset derived from the English GSM-Plus benchmark and verified by human translators. We evaluate six open-source LLMs Qwen3-32B, Llama-3.1-8B-Instant, Llama-3.3-70B-Versatile, Llama-4-Scout-17B-16E-Instruct, GPT-OSS-120B, and GPT-OSS-20B using a benchmark of 9,000 evaluation samples comprising 1,000 seed questions and 8,000 perturbed variants under both Standard Prompting and Chain-of-Thought (CoT) Prompting. Experimental results show that GPT-OSS-20B achieves the highest seed question accuracy of 96.08% under Standard Prompting, while larger models such as Llama-3.3-70B and GPT-OSS-120B demonstrate superior robustness across perturbation types. Furthermore, CoT prompting substantially improves reasoning for most models compared to Standard Prompting, yet a notable performance gap persists across all models relative to their English benchmarks, underscoring the inherent difficulty of perturbed Bengali text. This research makes a foundational contribution by providing GSM-PLUS-BN as a new resource and baseline for future Bengali mathematical reasoning research.

cs.CL