QuranicMMLU: A Cognitively-Aware Benchmark for Evaluating Generative AI Solutions on Quranic Linguistic Knowledge

TL;DR

QuranicMMLU benchmark evaluates generative AI on Quranic linguistic knowledge with an average multiple-choice accuracy of 84%.

cs.CL 🔴 Advanced 2026-09-19 28 views
Rawan El Ghali Umm Kulsoom Anas Madkoor Dima Faris Alsaudi Roaa Abdelmagid Roaa Ibrahim Raghad Mousa Hamza Aljaji Abdullah Khanafer Abdallah Alkanani Salah Feras Alali Rawan Khaled Mohamed Ehsaneddin Asgari
Generative AI Quran Linguistics Cognitive Evaluation NLP

Key Findings

Methodology

The study constructs a five-pillar Quranic taxonomy covering Phonology, Morphology, Syntax, Semantics, and Pragmatics. Questions are generated for each branch, stratified by Bloom's cognitive level, and independently answered and scored by LLMs, followed by manual review.

Key Results

  • Islamic-specialized models performed best in multiple-choice questions with an average accuracy of 96.4%, while open-ended questions averaged a quality score of 60%.
  • Multiple-choice accuracy averaged 84%, while open-ended scores were only 60%, indicating multiple-choice may obscure model deficiencies.
  • The rankings for multiple-choice and open-ended questions closely align (Kendall's τ=0.73), but multiple-choice may hide failures exposed in open-ended questions.

Significance

QuranicMMLU provides a rigorous, linguistically grounded framework for evaluating Arabic NLP capabilities in the Quranic domain. It addresses gaps in existing benchmarks by evaluating linguistic competency diversity and cognitive demand levels.

Technical Contribution

The study introduces a novel five-pillar taxonomy, combining multiple question formats and a multi-judge protocol to systematically evaluate knowledge across cognitive levels and verse perplexity dimensions.

Novelty

This is the first benchmark to stratify Quranic linguistic complexity, offering a more detailed evaluation of linguistic competencies than existing Islamic and Quranic QA resources.

Limitations

  • Open-ended responses are scored solely by LLM, which may introduce model bias.
  • Closed-source commercial models were not evaluated due to lack of API access.

Future Work

Future work includes expanding leaf coverage, adding human scoring of open-ended answers, and evaluating more Islamic-specialized systems.

AI Executive Summary

QuranicMMLU is a new benchmark for evaluating generative AI on Quranic Arabic. Existing benchmarks focus on general question answering and semantic retrieval without probing specific linguistic competencies or stratifying by cognitive demand and verse difficulty.

The study constructs a five-pillar Quranic taxonomy covering Phonology, Morphology, Syntax, Semantics, and Pragmatics. Questions are generated for each branch, stratified by Bloom's cognitive level and verse perplexity, then independently answered and scored by LLMs, followed by manual review.

Results show that Islamic-specialized models perform best in multiple-choice questions with an average accuracy of 96.4%, while open-ended questions average a quality score of 60%. This indicates that multiple-choice may obscure model deficiencies. QuranicMMLU provides a rigorous, linguistically grounded framework for evaluating Arabic NLP capabilities in the Quranic domain.

Deep Analysis

Background

The Quran is the primary religious and cultural text for nearly two billion Muslims, and engaging with it meaningfully demands not just reading but a genuine understanding of its language and content. As large language models are increasingly consulted on Islamic topics, accurate, fine-grained evaluation of their understanding of Quranic Arabic becomes crucial.

Core Problem

Existing Islamic and Quranic QA resources assess factual and reasoning ability but do not target the diversity of linguistic competencies required for Quranic understanding.

Innovation

QuranicMMLU introduces a five-pillar Quranic taxonomy covering Phonology, Morphology, Syntax, Semantics, and Pragmatics, providing a detailed framework for evaluating linguistic competencies.

Methodology

  • �� Construct a five-pillar Quranic taxonomy
  • �� Generate questions stratified by Bloom's cognitive level and verse perplexity
  • �� Use LLMs to independently answer and score
  • �� Conduct manual review and adjudication

Experiments

The study uses 980 human-reviewed questions covering linguistic competencies across five pillars. Each question is issued in both open-ended and multiple-choice form and benchmarked on 12 systems.

Results

Islamic-specialized models performed best in multiple-choice questions with an average accuracy of 96.4%, while open-ended questions averaged a quality score of 60%. Multiple-choice accuracy averaged 84%, indicating potential obscuration of model deficiencies.

Applications

QuranicMMLU provides a rigorous, linguistically grounded framework for evaluating Arabic NLP capabilities in the Quranic domain, suitable for researchers and developers.

Limitations & Outlook

Open-ended responses are scored solely by LLM, which may introduce model bias. Closed-source commercial models were not evaluated due to lack of API access.

Plain Language Accessible to non-experts

Imagine you're learning a new language with complex grammar and vocabulary rules. QuranicMMLU is like a detailed study guide that helps you understand these complex rules. It's not just a simple Q&A tool but a comprehensive evaluation system that tests your understanding across different linguistic levels.

ELI14 Explained like you're 14

Imagine you're playing a super complex game with lots of levels, each with different challenges. QuranicMMLU is like the game's guide, helping you pass each level step by step. It doesn't just give you answers but teaches you how to understand and solve problems. Isn't that cool?

Glossary

Phonology

The study of the sound system of a language, including pronunciation and phonemes.

Used to evaluate pronunciation rules in the Quran.

Morphology

The study of word formation and structure.

Used to analyze root and word form changes in the Quran.

Syntax

The study of sentence structure and grammatical rules.

Used to parse sentence structures in the Quran.

Semantics

The study of meaning and interpretation of words and sentences.

Used to understand word and sentence meanings in the Quran.

Pragmatics

The study of language use and contextual meaning.

Used to analyze context and communicative functions in the Quran.

Open Questions Unanswered questions from this research

  • 1 How to better evaluate the quality of open-ended answers, especially without human involvement.
  • 2 How to expand the benchmark to cover more linguistic phenomena and cognitive levels.

Applications

Immediate Applications

Linguistic Competency Evaluation

Researchers can use QuranicMMLU to evaluate generative AI performance on Quranic linguistic knowledge.

Long-term Vision

Educational Tool

QuranicMMLU could develop into an educational tool to help learners better understand Quranic language.

Abstract

We introduce QuranicMMLU, a benchmark for evaluating generative AI on Quranic Arabic across multiple dimensions of linguistic complexity. Existing Quranic benchmarks center on general question answering and semantic retrieval, without probing specific linguistic competencies or stratifying by cognitive demand and verse difficulty. We construct a five-pillar Quranic taxonomy spanning Phonology, Morphology, Syntax, Semantics, and Pragmatics, with 31 leaves covering phenomena from tajwīd and root-and-pattern morphology to occasions of revelation and inter-surah coherence. For each leaf we generate questions stratified by Bloom's cognitive level and verse perplexity, then have LLM as a judge to independently answer and score every item and route the annotations to manual review. The resulting dataset comprises 980 human-reviewed questions, each issued in both open-ended and multiple-choice form. We benchmark 12 systems on these items and find that the Islamic-specialized model leads, yet every system scores higher on multiple-choice accuracy (average 84%) than open-ended answer quality (average 60%): the two rankings agree closely (Kendall's τ=0.73), but multiple-choice scoring hides failures that surface only once answer choices are removed. QuranicMMLU thus offers a rigorous, linguistically grounded framework for evaluating Arabic NLP in the Quranic domain.

cs.CL