Khayyam Challenge (PersianMMLU): Is Your LLM Truly Wise to The Persian Language?

TL;DR

Khayyam Challenge evaluates Persian LLMs using 20,192 multiple-choice questions.

cs.CL 🔴 Advanced 2024-04-10 1 views
Omid Ghahroodi Marzia Nouri Mohammad Vali Sanian Alireza Sahebi Doratossadat Dastgheib Ehsaneddin Asgari Mahdieh Soleymani Baghshah Mohammad Hossein Rohban
Persian LLM evaluation multilingual education

Key Findings

Methodology

The Khayyam Challenge employs 20,192 four-choice questions from 38 tasks, covering literature, mathematics, and sciences, to assess LLMs' language comprehension and reasoning. The dataset includes rich metadata like difficulty levels and descriptive answers, ensuring comprehensive and accurate evaluation.

Key Results

  • GPT-4 outperformed all models across categories, with an average accuracy 8% higher than Claude3-haiku.
  • Aya performed comparably or better than GPT-3.5 in 8 tasks, showing the potential of open-source models.
  • PersianMind, despite being trained on Persian, underperformed compared to mT0XL.

Significance

The Khayyam Challenge provides a comprehensive evaluation framework for Persian LLMs, filling the gap in non-English language evaluation, advancing multilingual LLM development, and revealing limitations in technical disciplines.

Technical Contribution

By using original Persian data, the Khayyam Challenge avoids translation errors and preserves cultural nuances. Its scalable design allows for future data updates, supporting ongoing evaluations.

Novelty

This is the first comprehensive LLM evaluation framework specifically designed for Persian, combining multidisciplinary, multi-difficulty questions with rich metadata support.

Limitations

  • Some models perform poorly in technical fields, especially calculus and geometry.
  • The dataset is primarily based on educational exams and may not apply to all domains.

Future Work

Future work could expand the dataset's subject range and develop models better suited for Persian to improve performance in technical fields.

AI Executive Summary

The Khayyam Challenge (PersianMMLU) aims to evaluate large language models (LLMs) supporting the Persian language through 20,192 multiple-choice questions across 38 tasks, covering literature, mathematics, and sciences. The dataset's unique features include rich metadata such as difficulty levels and descriptive answers, ensuring comprehensive and accurate evaluation.

Experimental results show that GPT-4 outperformed all models across categories, with an average accuracy 8% higher than Claude3-haiku. However, Aya performed comparably or better than GPT-3.5 in multiple tasks, demonstrating the potential of open-source models. Despite being specifically trained for Persian, PersianMind underperformed compared to mT0XL.

The Khayyam Challenge provides a comprehensive evaluation framework for Persian LLMs, filling the gap in non-English language evaluation, advancing multilingual LLM development, and revealing limitations in technical disciplines. Future work could expand the dataset's subject range and develop models better suited for Persian to improve performance in technical fields.

Deep Analysis

Background

In recent years, large language models (LLMs) have made significant advances in machine intelligence applications. However, evaluations for non-English languages lag behind, resulting in a lack of robust LLM support for many languages. The Khayyam Challenge (PersianMMLU) addresses this by providing a comprehensive evaluation framework specifically for Persian.

Core Problem

Most existing LLM evaluation frameworks focus on English, lacking support for non-English languages. Persian, with its rich cultural background, requires a dedicated evaluation framework as direct translation methods often fall short.

Innovation

The Khayyam Challenge's innovation lies in its comprehensive coverage of 38 subjects, providing rich metadata support to ensure comprehensive and accurate evaluation. By using original Persian data, it avoids translation errors and preserves cultural nuances.

Methodology

  • �� The dataset is sourced from Persian exams, containing 20,192 four-choice questions.
  • �� Covers literature, mathematics, sciences, etc.
  • �� Provides rich metadata like difficulty levels and descriptive answers.
  • �� Uses original Persian data to avoid translation errors.

Experiments

Experiments evaluated multiple models, including GPT-4, GPT-3.5, and Aya, using zero-shot and chain-of-thought methods. Results showed GPT-4 performed best across all categories, while Aya excelled in several tasks.

Results

GPT-4 outperformed all models across categories, with an average accuracy 8% higher than Claude3-haiku. Aya performed comparably or better than GPT-3.5 in multiple tasks, demonstrating the potential of open-source models.

Applications

The Khayyam Challenge can be used to evaluate and improve Persian LLMs, advancing multilingual LLM development, particularly in educational and cultural-related fields.

Limitations & Outlook

Some models perform poorly in technical fields, especially calculus and geometry. The dataset is primarily based on educational exams and may not apply to all domains. Future work could expand the dataset's subject range.

Plain Language Accessible to non-experts

Imagine a school exam, the Khayyam Challenge is like a test specifically designed for Persian, with questions from various subjects like math, literature, and science. Each question has a detailed explanation to help students understand why the answer is correct. This exam not only tests students' knowledge but also helps teachers understand where students need more help. Like a comprehensive check-up, the Khayyam Challenge helps us understand the strengths and weaknesses of Persian LLMs.

ELI14 Explained like you're 14

Hey, imagine you're taking a super important school exam, but this time it's written in Persian! The Khayyam Challenge is like this exam, with lots of different questions from math to literature. Each question has a detailed explanation telling you why the answer is right. This challenge not only tests your knowledge but helps teachers understand where you need more help. It's like a full check-up, helping us understand the strengths and weaknesses of Persian LLMs.

Glossary

Large Language Model (LLM)

A large AI model capable of generating and understanding natural language.

Used to evaluate language comprehension in Persian.

Metadata

Data about data, such as difficulty levels and descriptive answers.

Provides additional information about questions.

Zero-shot method

An evaluation method that does not require training samples.

Used to assess model performance on new tasks.

Chain-of-thought

A step-by-step problem-solving method similar to human thinking.

Used to enhance model reasoning capabilities.

Persian

A language primarily spoken in Iran, with a rich cultural background.

The language specifically targeted by the Khayyam Challenge.

Open Questions Unanswered questions from this research

  • 1 How to improve LLM performance in technical fields? Current models perform poorly in calculus and geometry.
  • 2 How to expand the dataset's subject range to cover more domains?

Applications

Immediate Applications

Educational Assessment

The Khayyam Challenge can be used to assess students' Persian language abilities, helping teachers develop more effective teaching plans.

Long-term Vision

Multilingual LLM Development

Advancing multilingual LLM development, particularly in non-English language evaluation and application.

Abstract

Evaluating Large Language Models (LLMs) is challenging due to their generative nature, necessitating precise evaluation methodologies. Additionally, non-English LLM evaluation lags behind English, resulting in the absence or weakness of LLMs for many languages. In response to this necessity, we introduce Khayyam Challenge (also known as PersianMMLU), a meticulously curated collection comprising 20,192 four-choice questions sourced from 38 diverse tasks extracted from Persian examinations, spanning a wide spectrum of subjects, complexities, and ages. The primary objective of the Khayyam Challenge is to facilitate the rigorous evaluation of LLMs that support the Persian language. Distinctive features of the Khayyam Challenge are (i) its comprehensive coverage of various topics, including literary comprehension, mathematics, sciences, logic, intelligence testing, etc., aimed at assessing different facets of LLMs such as language comprehension, reasoning, and information retrieval across various educational stages, from lower primary school to upper secondary school (ii) its inclusion of rich metadata such as human response rates, difficulty levels, and descriptive answers (iii) its utilization of new data to avoid data contamination issues prevalent in existing frameworks (iv) its use of original, non-translated data tailored for Persian speakers, ensuring the framework is free from translation challenges and errors while encompassing cultural nuances (v) its inherent scalability for future data updates and evaluations without requiring special human effort. Previous works lacked an evaluation framework that combined all of these features into a single comprehensive benchmark. Furthermore, we evaluate a wide range of existing LLMs that support the Persian language, with statistical analyses and interpretations of their outputs.

cs.CL cs.AI