QFrBLiMP: a Quebec-French Benchmark of Linguistic Minimal Pairs

TL;DR

Introduces QFrBLiMP, a human-annotated Quebec-French minimal pairs benchmark, evaluating LLMs' grammatical competence with detailed analysis.

cs.CL 🔴 Advanced 2025-09-30 47 views
David Beauchemin Pier-Luc Veilleux Johanna-Pascale Roy Richard Khoury
NLP Language Models Linguistic Evaluation Dialectal Robustness Benchmarking

Key Findings

Methodology

Using 1,761 manually crafted minimal pairs across 20 grammatical phenomena, annotated by 12 native Quebec-French speakers, the study assesses LLMs’ grammatical acceptability via sentence probability comparisons. Models (77 open-source LLMs) are evaluated with perplexity scores, and statistical tests (Z-test) compare performance across dialects and model sizes. The approach combines fine-grained linguistic analysis with large-scale model benchmarking, providing insights into the correlation between model size and grammatical competence, as well as limitations in deep semantic understanding.

Key Results

  • Model performance correlates positively with size, with the largest models achieving over 90% accuracy on frequent phenomena, but all models underperform on deep semantic phenomena, lagging behind human accuracy (~93%).
  • Compared to MultiBLiMP-Fr, Quebec-French models show significant performance drops, yet top models remain within statistical significance bounds, indicating cross-dialectal robustness.
  • Frequent, rule-based phenomena are well learned (>95%), while complex semantic and long-distance dependencies reveal persistent weaknesses, highlighting the gap in deep semantic comprehension.

Significance

This work pioneers a high-quality, human-annotated benchmark for Quebec-French, addressing the gap in dialect-specific NLP evaluation. It demonstrates that larger models improve grammatical accuracy but still struggle with deep semantics, guiding future research in multilingual and dialectal NLP. The dataset’s real-world relevance and detailed analysis provide a foundation for targeted model improvements, fostering progress toward more linguistically aware AI systems.

Technical Contribution

The paper introduces QFrBLiMP, a linguistically grounded, human-annotated benchmark for Quebec-French, combined with statistical evaluation methods. It offers a detailed, category-wise performance analysis across 77 models, revealing the relationship between model size and grammatical competence, and exposing persistent semantic limitations. This approach advances the state-of-the-art in dialect-specific NLP evaluation, emphasizing the importance of real data and nuanced analysis.

Novelty

First to create a high-quality, human-annotated Quebec-French grammatical benchmark based on official linguistic sources. Unlike synthetic datasets, QFrBLiMP ensures linguistic authenticity and comprehensive coverage of 20 phenomena. It extends the scope of multilingual benchmarks, emphasizing dialectal variation, and offers a new paradigm for evaluating NLP models in less-resourced languages and dialects.

Limitations

  • Data is sourced from formal linguistic resources, which may not fully capture colloquial or informal usage prevalent in everyday speech, limiting real-world applicability.
  • Evaluation relies primarily on perplexity, which may not fully reflect deep semantic understanding or contextual reasoning capabilities.
  • Model training data may not encompass the full diversity of Quebec French, suggesting future work should include more varied and spontaneous speech corpora.

Future Work

Future directions include expanding the corpus with colloquial and spoken language data, adding more complex semantic and pragmatic phenomena, and integrating multimodal information. Developing more comprehensive metrics beyond perplexity, such as semantic similarity or entailment, will better assess deep understanding. Additionally, extending evaluation to other dialects and low-resource languages can foster more inclusive NLP systems.

AI Executive Summary

The rapid advancement of large language models (LLMs) has revolutionized natural language processing, yet their linguistic understanding across diverse dialects remains underexplored. While benchmarks like GLUE and BLiMP have evaluated models mainly on English and standard French, dialectal varieties such as Quebec French lack dedicated assessment tools. This gap limits our understanding of how well models grasp dialect-specific grammatical phenomena.

Addressing this, the authors introduce QFrBLiMP—a high-quality, human-annotated benchmark comprising 1,761 minimal pairs across 20 grammatical phenomena, specifically targeting Quebec French. Extracted from official linguistic resources, each sentence pair was annotated by 12 native speakers, ensuring linguistic authenticity. The benchmark evaluates 77 open-source models using perplexity scores, comparing the probability favorability of grammatical versus ungrammatical sentences.

Results reveal a clear scale effect: larger models outperform smaller ones, with accuracy reaching over 90% on frequent phenomena. However, all models significantly underperform on phenomena requiring deep semantic understanding, lagging behind human accuracy (~93%). When compared to the broader MultiBLiMP-Fr benchmark, Quebec French models show performance drops, yet top models maintain statistical robustness, indicating cross-dialectal resilience.

These findings underscore that model scaling improves grammatical competence but does not fully bridge the gap in semantic comprehension. The study emphasizes the importance of authentic, dialect-specific evaluation data and provides a pathway for future research to enhance deep understanding, extend to informal speech, and develop more nuanced metrics. Ultimately, this work advances the development of linguistically aware AI capable of handling diverse dialectal varieties, fostering more inclusive NLP applications.

Deep Analysis

Background

随着大规模预训练语言模型(LLMs)在NLP中的广泛应用,评估其语法和语义能力成为研究重点。早期工作如GLUE、BLiMP等主要集中在英语和标准法语,缺乏对少数方言的系统性评估。近年来,基于最小对比对(MPs)的方法逐渐兴起,利用真实语料和人类标注,提升评估的真实性和细粒度。多语种基准如CLiMP、JBLiMP、MultiBLiMP已在多语言环境中展开,但针对魁北克法语等少见变体的研究仍不足。魁北克法语具有独特的语法特征和用法,亟需专门的评估工具来衡量模型的理解能力。

Core Problem

现有模型在英语和主流法语中的表现虽有提升,但在魁北克法语等少数方言中的语法理解仍不充分。缺乏高质量、真实语料的评估基准,导致模型在实际应用中可能出现偏差。尤其是在深层语义理解和复杂句法结构方面,模型表现仍远低于人类水平。如何构建具有代表性、覆盖丰富语法现象的评估体系,成为亟待解决的核心问题。此外,模型在处理方言变体时的鲁棒性和泛化能力仍待提升。

Innovation

本研究的创新点包括:1)基于官方“Banque de dépannage linguistique”语料库,构建高质量的魁北克法语最小对比语料库QFrBLiMP,确保语料的真实性和规范性;2)涵盖20个语法类别,系统评估模型在不同语法层级的能力;3)采用人类多 annotator标注,确保标注质量,结合统计检验分析模型在不同语法现象和模型规模上的表现差异。这些创新为少数方言的模型评估提供了新范例,也丰富了多语种语法认知研究的内容。

Methodology

  • �� 数据采集:从官方“Banque de dépannage linguistique”中手工提取句子,标注语法正确性,组织成20类语法现象的最小对比对。
  • �� 人类标注:由12名魁北克法语母语者使用定制工具进行标注,确保一致性,采用多数投票确定最终标签。
  • �� 模型评估:利用perplexity指标对77个开源模型进行句子概率评分,比较模型对句子正确性偏好的差异。
  • �� 统计分析:采用Z检验比较模型在魁北克法语与多语种基准中的表现差异,分析模型规模与能力的关系。

Experiments

实验设计包括:选择参数规模跨度大的模型,使用perplexity作为核心指标,结合人类标注作为参考基准。评估指标为准确率,分析模型在不同语法类别上的表现差异。通过统计检验验证模型在两个基准上的性能差异,确保结果具有统计显著性。还进行了模型规模与性能的相关性分析,揭示模型规模与语法能力的关系,为模型优化提供依据。

Results

模型规模越大,语法准确率越高,最高超过90%。在规则性强、频繁出现的语法现象上表现优异,达95%以上,但在深层语义和复杂依存关系上表现明显不足,所有模型都低于人类水平(约93%)。魁北克法语模型在整体表现上低于多语种基准,但最优模型仍在统计显著区间内,显示出一定的跨方言鲁棒性。这些结果表明,模型在规模扩大后,能力逐步提升,但深层理解仍是瓶颈。

Applications

该基准可用于评估多语种模型在少数方言中的语法能力,为模型优化和多语种适应提供指标。未来可结合实际应用场景,如智能校对、语音识别、对话系统等,提升模型在本地化语料中的表现,推动多语种NLP的应用发展。

Limitations & Outlook

数据主要来源于官方规范语料,可能偏向正式书面语,未充分覆盖口语和非正式用法。评估指标主要依赖perplexity,未结合其他语义理解指标,限制对深层语义能力的全面评估。模型训练数据有限,未来需引入更多多样化、口语化的语料,以提升模型的泛化能力和适应性。

Plain Language Accessible to non-experts

想象你在一家工厂里,工人们每天都在按照一套规则组装产品。不同的工人可能擅长不同的任务,有的熟悉拼装,有的擅长检查。大模型就像这些工人,学会了很多规则,但在面对复杂的订单时,可能会出错,比如理解深层的指令或处理特殊情况。这个研究就像考核工人们是否真正掌握了所有规则,特别是在魁北克法语这种特殊“工厂语言”中。通过让工人们完成不同的拼装任务,研究者可以评估他们的技能水平,找出哪些规则还需要加强。最终,帮助工厂的工人变得更聪明,更能应对各种复杂的订单。

ELI14 Explained like you're 14

想象你在学校学语法,你知道句子要符合规则才能算正确。有时候一句话虽然看起来不错,但其实不符合语法。科学家们也遇到这个问题,他们想知道电脑程序是不是也能像人一样理解这些规则。于是,他们设计了一种考试,给电脑一些句子,让它判断哪个正确,哪个错。这个考试就像是测验,题目是两个句子,一个符合语法,一个不符合。科学家用很多真实的句子,特别是来自魁北克的法语,来测试电脑的“语法水平”。他们发现,越大的模型越像人,能正确判断的句子也越多,但在理解深层意思和复杂句子方面,还是比不过人类。这就像你学会了很多规则,但在理解复杂故事时还得努力。这个研究帮助我们知道,未来的电脑会变得更聪明,也提醒我们要继续改进它们的理解能力。

Abstract

In this paper, we introduce the Quebec-French Benchmark of Linguistic Minimal Pairs (QFrBLiMP), a corpus designed to evaluate LLMs' linguistic knowledge of prominent grammatical phenomena in Quebec-French. QFrBLiMP comprises 1,761 minimal pairs annotated with 20 LPs. Specifically, these minimal pairs have been created by manually modifying sentences extracted from an official online resource maintained by a Québec government institution. Each pair is annotated by 12 Quebec-French native speakers, who select the sentence they consider grammatical from the two. These annotations are used to compare the competency of LLMs with that of humans. We evaluate different LLMs on QFrBLiMP and MultiBLiMP-Fr by observing the rate of higher probabilities assigned to the sentences of each minimal pair for each category. We find that while grammatical competence scales with model size, a clear hierarchy of difficulty emerges. All benchmarked models consistently fail on phenomena requiring deep semantic understanding, revealing a critical limitation. Finally, our statistical analysis comparing QFrBLiMP and MultiBLiMP reveals a significant performance degradation for most models on Quebec-French; however, the most capable models remain within the statistical significance interval, demonstrating cross-dialectal robustness.

cs.CL