The Interspeech 2026 Audio Reasoning Challenge: Evaluating Reasoning Process Quality for Audio Reasoning Models and Agents

TL;DR

The Interspeech 2026 Audio Reasoning Challenge uses MMAR-Rubrics to evaluate audio reasoning models, with agent systems leading in quality.

cs.SD 🔴 Advanced 2026-02-16 45 views
Ziyang Ma Ruiyang Xu Yinghao Ma Chao-Han Huck Yang Bohan Li Jaeyeon Kim Jin Xu Jinyu Li Carlos Busso Kai Yu Eng Siong Chng Xie Chen
audio reasoning large language models multimodal reinforcement learning explainable AI

Key Findings

Methodology

The study introduces MMAR-Rubrics, focusing on the factuality and logic of reasoning chains. The challenge features Single Model and Agent tracks, evaluating end-to-end models and multimodal agents. Gemini-2.5-Pro generates verifiable criteria, with GPT-4o assessing reasoning paths.

Key Results

  • Agent systems lead in reasoning quality with a top score of 69.83%, while the best single model scores 65.29%.
  • Single models rapidly advance through reinforcement learning and data pipelines, using GRPO to optimize reasoning.
  • Agent systems achieve higher reasoning transparency through tool integration and cross-modal analysis.

Significance

This challenge shifts audio intelligence evaluation from result-oriented to process-oriented metrics. By introducing instance-level evaluation protocols, it significantly improves the reliability and stability of reasoning quality, offering new insights for explainable audio intelligence.

Technical Contribution

This is the first challenge dedicated to evaluating the quality of reasoning processes in the audio domain, introducing MMAR-Rubrics for instance-level evaluation, addressing instability issues in previous methods. The dual-track design reveals the strengths of both end-to-end models and agent systems.

Novelty

This is the first challenge focusing on audio reasoning process quality, achieving finer-grained evaluation through MMAR-Rubrics, significantly different from previous benchmarks that only focused on final answers.

Limitations

  • The current evaluation protocol relies on manually annotated reasoning paths, which may introduce subjective bias.
  • Single models still have limitations in complex reasoning tasks and need further optimization.

Future Work

Future work will explore more automated reasoning evaluation methods, improve model performance in multi-step reasoning tasks, and extend to more audio scenarios.

AI Executive Summary

Audio reasoning is a crucial aspect of human intelligence, yet existing large audio language models lack transparency in reasoning. The Interspeech 2026 Audio Reasoning Challenge, using the MMAR-Rubrics evaluation protocol, is the first to focus on the quality of audio reasoning processes. The challenge attracted 156 teams from 18 countries, divided into Single Model and Agent tracks. Results show that agent systems excel in reasoning quality, benefiting from tool integration and cross-modal analysis, while single models rapidly advance through reinforcement learning and data pipelines. This challenge provides new insights for explainable audio intelligence, shifting evaluation from result-oriented to process-oriented metrics.

The MMAR-Rubrics protocol significantly improves the reliability and stability of reasoning quality through instance-level evaluation. Agent systems achieve higher reasoning transparency through tool integration and cross-modal analysis, while single models optimize reasoning processes using GRPO. The main contributions include the first introduction of reasoning quality evaluation in the audio domain, proposing MMAR-Rubrics for instance-level evaluation, addressing instability issues in previous methods.

Despite significant progress, the current evaluation protocol relies on manually annotated reasoning paths, which may introduce subjective bias. Future work will explore more automated reasoning evaluation methods, improve model performance in multi-step reasoning tasks, and extend to more audio scenarios.

Deep Analysis

Background

Audio reasoning is a key research area in artificial intelligence. Recent advances in large language models and audio processing have led to significant progress in audio reasoning models. However, existing models still lack transparency, especially in complex multi-step reasoning tasks. Previous benchmarks mainly focused on the accuracy of final answers, neglecting the quality of intermediate reasoning processes, posing risks in real-world applications.

Core Problem

Existing audio reasoning models lack transparency and stability, especially in complex multi-step reasoning tasks. Traditional evaluation methods focus on final answer accuracy, neglecting the quality of intermediate reasoning processes, posing risks in real-world applications.

Innovation

This study introduces the first reasoning quality evaluation in the audio domain, proposing the MMAR-Rubrics instance-level evaluation protocol. This protocol significantly improves the reliability and stability of reasoning quality through instance-level evaluation. The dual-track design reveals the strengths of both end-to-end models and agent systems, offering new insights for explainable audio intelligence.

Methodology

  • �� Introduce MMAR-Rubrics evaluation protocol, focusing on the factuality and logic of reasoning chains.
  • �� The challenge features Single Model and Agent tracks, evaluating end-to-end models and multimodal agents.
  • �� Use Gemini-2.5-Pro to generate verifiable criteria, with GPT-4o assessing reasoning paths.

Experiments

The experimental design includes two stages: a preliminary stage with a subset of 500 questions and a final stage with the complete MMAR benchmark of 1000 questions. The Single Model track requires end-to-end models to perform reasoning in a single forward pass, while the Agent track allows the use of multiple tools and models in collaboration.

Results

Agent systems lead in reasoning quality with a top score of 69.83%, while the best single model scores 65.29%. Single models rapidly advance through reinforcement learning and data pipelines, using GRPO to optimize reasoning. Agent systems achieve higher reasoning transparency through tool integration and cross-modal analysis.

Applications

Applications include intelligent voice assistants, audio analysis systems, and multimodal interaction platforms. Enhancing transparency and stability in audio reasoning models can significantly improve performance in complex scenarios.

Limitations & Outlook

The current evaluation protocol relies on manually annotated reasoning paths, which may introduce subjective bias. Single models still have limitations in complex reasoning tasks and need further optimization. Future work will explore more automated reasoning evaluation methods, improve model performance in multi-step reasoning tasks, and extend to more audio scenarios.

Plain Language Accessible to non-experts

Imagine you're in a kitchen cooking a meal. Audio reasoning is like creating a delicious dish based on ingredients and cooking steps. Existing audio models are like chefs who only focus on the final dish, ignoring the cooking process. The Interspeech 2026 challenge is like a new cooking competition that requires not only a tasty dish but also a demonstration of every cooking step. The MMAR-Rubrics evaluation protocol is like a strict judge focusing on each detail, ensuring every step is logical and reasonable. This way, we can better understand the reasoning process of audio models and improve their performance in complex scenarios.

ELI14 Explained like you're 14

Hey there! Imagine you're playing a mystery game. Audio reasoning is like solving puzzles based on sound clues. Existing audio models are like players who only focus on the final answer, ignoring the reasoning process. The Interspeech 2026 challenge is like a new competition that requires not only solving the puzzle but also showing every reasoning step. The MMAR-Rubrics evaluation protocol is like a strict referee focusing on each detail, ensuring every step is logical and reasonable. This way, we can better understand the reasoning process of audio models and improve their performance in complex scenarios. Cool, right?

Glossary

MMAR-Rubrics

An instance-level evaluation protocol focusing on the factuality and logic of reasoning chains.

Used to evaluate the reasoning quality of audio reasoning models.

Agent System

A multimodal system that uses multiple tools and models in collaboration for reasoning.

Performed best in the challenge, leading in reasoning quality.

Reinforcement Learning

A machine learning method that optimizes model behavior through reward signals.

Used to optimize the reasoning process of single models.

GRPO

A reinforcement learning algorithm used to optimize the reasoning process of models.

Used in the Single Model track to enhance reasoning quality.

Cross-Modal Analysis

An analysis method combining multiple data sources.

Agent systems achieve higher reasoning transparency through cross-modal analysis.

Open Questions Unanswered questions from this research

  • 1 How to achieve automated reasoning evaluation without relying on manual annotations?
  • 2 How to improve single model performance in complex reasoning tasks?

Applications

Immediate Applications

Intelligent Voice Assistants

Enhancing reasoning transparency improves performance in complex scenarios.

Long-term Vision

Multimodal Interaction Platforms

Enhancing intelligence through cross-modal analysis.

Abstract

Recent Large Audio Language Models (LALMs) excel in understanding but often lack transparent reasoning. To address this "black-box" limitation, we organized the Audio Reasoning Challenge at Interspeech 2026, the first shared task dedicated to evaluating Chain-of-Thought (CoT) quality in the audio domain. The challenge introduced MMAR-Rubrics, a novel instance-level protocol assessing the factuality and logic of reasoning chains. Featured Single Model and Agent tracks, the competition attracting 156 teams from 18 countries and regions. Results show agent systems currently lead in reasoning quality, utilizing iterative tool orchestration and cross-modal analysis. Besides, single models are rapidly advancing via reinforcement learning and sophisticated data pipeline. We details the challenge design, methodology, and a comprehensive analysis of state-of-the-art systems, providing new insights for explainable audio intelligence.

cs.SD cs.CL cs.MM