UrBLiMP: A Benchmark for Evaluating the Linguistic Competence of Large Language Models in Urdu

TL;DR

Introduces UrBLiMP, a benchmark with 5696 minimal pairs for evaluating Urdu syntax in multilingual LLMs, achieving up to 94.73% accuracy.

cs.CL 🔴 Advanced 2025-08-02 49 views
Farah Adeeba Brian Dillon Hassan Sajjad Rajesh Bhatt
multilingual evaluation low-resource languages syntactic understanding language benchmark LLM assessment

Key Findings

Methodology

UrBLiMP was constructed through a hybrid approach combining Urdu Treebank templates and surface pattern matching on diverse corpora, followed by manual validation. Covering 10 syntactic phenomena, it includes 5696 minimal pairs designed to test grammatical acceptability. Human annotation yielded 96.10% agreement, ensuring data reliability. Twenty multilingual models were evaluated using perplexity and accuracy metrics, revealing significant performance variation across phenomena, with LLaMA-3-70B reaching 94.73% average accuracy. The dataset enables fine-grained analysis of models’ syntactic knowledge in Urdu.

Key Results

  • LLaMA-3-70B achieved the highest average accuracy of 94.73%, with models like Gemma-3-27B-PT close behind. Performance varied notably across phenomena; for example, Aspect Agreement and Ergativity were well captured, but long-distance dependencies remained challenging. Fine-tuned models like Alif-1.0-8B-Instruct improved results, highlighting the importance of continued pretraining. Human performance was 96.10%, validating dataset quality. The results underscore both the potential and current limitations of multilingual LLMs in low-resource languages.
  • Model performance differed significantly across phenomena, with some models surpassing LLaMA-3-70B on specific tasks, indicating specialized strengths. The evaluation revealed persistent challenges in modeling long-range syntactic dependencies and morphological variations, especially in morphologically rich Urdu. Fine-tuning improved performance in certain areas, demonstrating the benefit of targeted adaptation. The dataset’s detailed phenotypic coverage allows precise diagnosis of model weaknesses, guiding future development.
  • The high human agreement rate (96.10%) confirms UrBLiMP’s reliability as a benchmark. The performance gap between models and humans highlights ongoing challenges in capturing nuanced syntactic rules. The dataset’s design and evaluation framework set a new standard for low-resource language assessment, fostering more linguistically informed model development. Overall, the study advances understanding of multilingual models’ capabilities in morphologically complex, low-resource settings, providing a foundation for future research.

Significance

This work addresses a critical gap in NLP evaluation for low-resource languages, providing a systematic, linguistically grounded benchmark for Urdu. By enabling detailed analysis of models’ syntactic understanding, it facilitates targeted improvements, promoting fairness and inclusivity in multilingual NLP. The benchmark’s comprehensive coverage of core phenomena helps identify specific weaknesses, guiding model development and fine-tuning strategies. It also establishes a template for extending similar evaluation frameworks to other low-resource languages, fostering broader linguistic diversity in AI. The results inform both academia and industry, especially in applications like machine translation, grammar correction, and language learning tools, where syntactic accuracy is paramount. Ultimately, UrBLiMP contributes to building more linguistically aware and equitable AI systems.

Technical Contribution

The study introduces a novel hybrid data generation approach combining treebank templates and surface pattern matching, ensuring high-quality, diverse minimal pairs across ten key phenomena. It develops a comprehensive evaluation framework employing perplexity and accuracy metrics tailored for multilingual models, enabling fine-grained syntactic assessment in low-resource languages. The methodology includes rigorous manual validation, ensuring dataset reliability. The evaluation of 20 models, from encoder-only to decoder-only architectures, provides empirical insights into their syntactic capabilities, especially in morphologically complex Urdu. This work advances the state-of-the-art in low-resource language evaluation, integrating linguistic theory with scalable data construction and model assessment techniques.

Novelty

UrBLiMP is the first systematic, phenotypically comprehensive benchmark for Urdu syntax, combining hybrid data generation with detailed linguistic phenomena coverage. Unlike existing resources limited to high-resource languages or small-scale tests, it offers a large, validated dataset specifically targeting the intricacies of Urdu grammar. Its methodology can be adapted to other low-resource languages, making it a versatile framework. The evaluation of diverse models highlights the current state and gaps in multilingual NLP, emphasizing the need for targeted pretraining and fine-tuning. This work bridges linguistic theory and practical model assessment, representing a significant step forward in low-resource NLP research.

Limitations

  • Despite high-quality manual validation, the dataset may not fully capture the natural variability of spontaneous Urdu speech and writing, potentially limiting ecological validity.
  • Models still struggle with long-distance dependencies and morphological complexity, indicating that current architectures lack robust syntactic generalization in morphologically rich languages.
  • Evaluation focuses primarily on grammatical acceptability, without integrating semantic or contextual factors, which are also crucial for real-world language understanding.

Future Work

未来将扩大UrBLiMP的语法现象覆盖范围,增加多样化的句法结构和语料来源,提升数据的代表性。结合深度学习与规则方法,增强模型对长距离依存和形态变化的理解能力。同时,计划引入跨模态评估,结合语义和上下文信息,推动低资源语种的多维理解研究。还将探索模型微调策略,提升在低资源环境中的泛化能力,促进多语种模型的公平性和实用性。

AI Executive Summary

UrBLiMP represents the first comprehensive syntactic benchmark tailored for Urdu, a low-resource language with rich morphology. Comprising 5696 minimal pairs across ten core phenomena, it was meticulously constructed through a hybrid approach combining treebank templates and surface pattern matching, followed by rigorous manual validation. This dataset enables detailed evaluation of multilingual large language models (LLMs), revealing significant performance variation. The top-performing model, LLaMA-3-70B, achieved an average accuracy of 94.73%, yet still exhibited notable weaknesses in long-distance dependencies and morphological complexity. These findings highlight both the potential of current models and their limitations in capturing fine-grained syntactic knowledge in low-resource languages. The study underscores the importance of continued pretraining and targeted fine-tuning, especially for morphologically rich languages like Urdu. By providing a linguistically grounded evaluation framework, UrBLiMP facilitates more precise diagnostics and improvements in multilingual NLP, fostering equitable AI development. The dataset’s high human agreement (96.10%) affirms its reliability, serving as a benchmark for future research. Overall, this work advances the understanding of multilingual models’ capabilities in low-resource settings, paving the way for more linguistically informed and inclusive AI systems.

Deep Analysis

Background

近年来,大型多语种预训练模型(如mT5、BLOOM、LLaMA)在自然语言处理领域取得了突破,但在低资源语言如乌尔都语中的表现仍有限。现有评估资源多集中于英语或汉语,缺乏针对乌尔都语复杂句法结构的细粒度测试。Warstadt等(2020)提出的BLiMP为英语提供了丰富的句法最小对句资源,但乌尔都语缺乏类似的系统评估工具。近年来,RuBLiMP和Jumelet等(2025)虽有所发展,但规模不足,难以全面衡量模型在乌尔都语中的句法理解能力。本研究旨在填补这一空白,构建系统化的乌尔都语句法评估基准,推动低资源语种的模型发展。

Core Problem

当前多语种大模型在乌尔都语中的句法理解能力不足,尤其在长距离依存、形态变化和复杂句法结构方面表现有限。这限制了模型在机器翻译、语法校正等实际应用中的效果。缺乏系统化的评估工具,使得模型的语法能力难以量化和比较,阻碍了模型的持续优化。设计高质量、多现象覆盖的评估数据集成为亟待解决的问题。

Innovation

本研究的创新点包括:1)结合乌尔都树库模板与大规模语料的表面模式,构建涵盖10个核心句法现象的最小对句数据集;2)引入多语种大模型在低资源环境中的系统评估框架,采用伪困惑度和准确率指标,量化模型的句法理解能力;3)通过人工验证确保数据质量,提供高可靠性评估工具。这些创新推动了低资源语种句法能力的系统评估,为模型优化提供实证依据。

Methodology

  • �� 采集乌尔都树库(Urdu Treebank)和多样化语料库作为数据源。
  • �� 设计针对10个句法现象的转换规则,生成最小对句。
  • �� 利用正则表达式和模式匹配,从树库和大语料中提取候选句。
  • �� 进行人工验证,确保句子符合目标句法结构。
  • �� 计算模型在接受度上的差异,采用伪困惑度和准确率指标。
  • �� 统计分析模型表现,比较不同模型和参数规模的差异,验证能力。

Experiments

采用20个多语种模型,包括LLaMA、Gemma、BERT、mT5等,参数范围从110M到70B。评估指标为伪困惑度和准确率,衡量模型对正确与错误句子的偏好。模型在不同句法现象上的表现差异被详细分析,特别关注长距离依存和形态变化。通过交叉验证和统计检验,确保结果稳健。还比较微调模型(如Alif-1.0-8B-Instruct)与预训练模型的性能差异,验证微调效果。

Results

最高模型LLaMA-3-70B达94.73%的平均准确率,但在长距离依存和复杂句法结构上仍有限制。模型在Aspect Agreement和Ergativity等现象表现优异,但在长距离关系和形态变化方面存在明显不足。微调模型如Alif-1.0-8B-Instruct在部分现象上优于基础模型。模型性能在不同句法类别间差异显著,揭示模型对细粒度句法结构的理解仍需提升。整体验证了UrBLiMP的有效性,为未来优化提供方向。

Applications

该基准可用于评估和优化多语种模型在乌尔都语中的句法理解能力,推动低资源语言的机器翻译、语法校正和自动文本生成等应用。结合语义和上下文信息,未来可提升模型的整体理解能力,满足实际需求。

Limitations & Outlook

模型在长距离依存和复杂句法结构上表现不足,主要因模型对形态变化和跨句关系理解有限。数据生成依赖模板,可能未涵盖所有自然语料的多样性。评估仅关注句法接受度,未结合语义和上下文,未来应结合多模态信息,提升模型整体理解。

Plain Language Accessible to non-experts

想象你在一家工厂工作,工厂里有许多不同的机器,每台机器负责不同任务。大模型就像这些机器,它们可以帮你写文章、翻译语言,但在理解复杂句子时,有时会出错,就像某些机器不能很好理解长长的生产线。为了让这些机器更聪明,科学家设计了一套测试题,就像给机器出一道题,看看它是否能正确判断一句话是否符合语法规则。通过不断改进这些测试和机器,最终希望让它们像人一样理解语言的细节。这项工作就是在做这样的测试,特别是针对乌尔都语这种资源较少的语言,帮助机器学会更好地理解它的句法结构。

ELI14 Explained like you're 14

想象你在学校里,有一台超级聪明的机器人老师,它可以帮你写作文、翻译,但有时会搞错一些语法规则。科学家们想让它变得更聪明,所以设计了一些特别的练习题,专门测试它是否懂得句子里的语法,比如词序、动词和主语的配合。这些练习题就像出一道难题,看看它能不能正确区分正确和错误的句子。研究人员用很多真实的乌尔都语句子,确保题目贴近实际。最后发现,虽然机器人在一些简单的语法上表现很好,但在处理长距离关系或复杂句子时还会出错。这项研究帮助我们知道,未来要让机器人更懂语言,还需要不断改进它的理解能力,就像我们不断学习一样。

Abstract

Multilingual Large Language Models (LLMs) have shown remarkable performance across various languages; however, they often include significantly less data for low-resource languages such as Urdu compared to high-resource languages like English. To assess the linguistic knowledge of LLMs in Urdu, we present the Urdu Benchmark of Linguistic Minimal Pairs (UrBLiMP) i.e. pairs of minimally different sentences that contrast in grammatical acceptability. UrBLiMP comprises 5,696 minimal pairs targeting ten core syntactic phenomena, carefully curated using the Urdu Treebank and diverse Urdu text corpora. A human evaluation of UrBLiMP annotations yielded a 96.10% inter-annotator agreement, confirming the reliability of the dataset. We evaluate twenty multilingual LLMs on UrBLiMP, revealing significant variation in performance across linguistic phenomena. While LLaMA-3-70B achieves the highest average accuracy (94.73%), its performance is statistically comparable to other top models such as Gemma-3-27B-PT. These findings highlight both the potential and the limitations of current multilingual LLMs in capturing fine-grained syntactic knowledge in low-resource languages.

cs.CL