SVHalluc: Benchmarking Speech-Vision Hallucination in Audio-Visual Large Language Models

TL;DR

Proposed SVHalluc benchmark evaluates speech-vision hallucination, revealing poor cross-modal alignment in state-of-the-art models with near-random accuracy.

eess.AS 🔴 Advanced 2026-05-31 78 views
Chenshuang Zhang Kyeong Seon Kim Chengxin Liu Tae-Hyun Oh
multimodal understanding hallucination detection speech-vision alignment benchmarking deep learning

Key Findings

Methodology

This work introduces SVHalluc, a comprehensive benchmark combining six tasks across semantic and temporal dimensions, utilizing the YouCook2 dataset. Automated sample generation with human verification creates positive and negative pairs for evaluating models like Qwen3-Omni and Gemini 2.5 Pro. Metrics such as accuracy and F1 are used to diagnose hallucination modes. The benchmark reveals that open-source models perform near chance levels (~50%) on most tasks, while Gemini 2.5 Pro achieves over 93% accuracy in global semantic alignment, highlighting significant gaps in cross-modal understanding.

Key Results

  • Open-source models in global semantic alignment hover around 50% accuracy, indicating random guess levels, whereas Gemini 2.5 Pro reaches 93.10%, demonstrating superior semantic grounding.
  • In fine-grained semantic tasks, models like Qwen2.5-Omni and Qwen3-Omni outperform others but still lag behind Gemini, showing limited ability to identify specific objects and actions.
  • Temporal tasks reveal that most models perform poorly, with accuracy near chance, confirming weak temporal reasoning. Gemini excels in temporal understanding, validating the benchmark's effectiveness in exposing these deficiencies.

Significance

This study exposes fundamental limitations of current multimodal large language models in aligning speech content with visual scenes, especially regarding complex semantics and temporal relationships. It addresses critical safety and reliability issues, guiding future research toward models with robust cross-modal reasoning. The benchmark provides a standardized, detailed evaluation framework, fostering progress in video understanding, virtual assistants, and multimedia AI applications. By systematically diagnosing hallucination modes, it helps identify core bottlenecks and directs efforts to improve model trustworthiness and interpretability in real-world deployments.

Technical Contribution

The paper introduces SVHalluc, the first benchmark targeting speech-vision hallucination through multi-level tasks. It combines automated data synthesis with human validation, covering global and fine-grained semantic alignment, cross-modal event binding, and temporal reasoning. The framework enables precise diagnosis of model failures, promoting targeted improvements. This approach advances the state-of-the-art by providing a comprehensive, scalable, and rigorous evaluation platform for multimodal AI, emphasizing the importance of cross-modal understanding and temporal reasoning in reducing hallucinations.

Novelty

This is the first systematic benchmark explicitly designed to evaluate speech-vision hallucination across semantic and temporal dimensions. Unlike prior benchmarks focusing solely on environmental sounds, SVHalluc emphasizes the rich semantics and complex time relationships inherent in speech, uncovering new failure modes. Its multi-task design and automated data pipeline set a new standard for comprehensive multimodal evaluation, pushing the boundaries of current understanding and fostering development of more reliable models.

Limitations

  • The benchmark relies on static question-answering tasks, lacking multi-turn dialogue or generative scenarios that are common in real-world applications.
  • Data diversity is limited to YouCook2, which may not fully capture the complexity of real-world videos and speech variations.
  • Evaluation focuses on accuracy metrics without considering model interpretability or robustness under adversarial conditions, suggesting future work should incorporate these aspects.

Future Work

Future directions include expanding to multi-turn dialogues and generative tasks, integrating larger and more diverse datasets, and exploring model architectures that better fuse semantic and temporal information. Additionally, developing explainability tools and robustness assessments will be crucial to enhance model trustworthiness. The authors also plan to investigate how training strategies can mitigate hallucinations, aiming for models capable of more accurate, context-aware multimodal reasoning in complex scenarios.

AI Executive Summary

In recent years, multimodal large language models have made significant strides in understanding complex visual and auditory data. However, despite their impressive capabilities, these models often suffer from hallucinations—producing plausible but ungrounded outputs that undermine reliability. This issue becomes particularly critical when models attempt to align speech content with visual scenes, where rich semantics and intricate temporal relationships pose substantial challenges. Existing benchmarks primarily focus on environmental sounds as indicators of events, leaving the more complex domain of human speech largely unexplored.

To address this gap, the authors propose SVHalluc, a comprehensive benchmark designed explicitly to evaluate speech-vision hallucination across two core dimensions: semantic and temporal. The benchmark comprises six carefully crafted tasks, ranging from global semantic alignment to fine-grained object recognition and temporal reasoning, all built upon the YouCook2 dataset. These tasks are formulated as question-answering prompts, enabling systematic diagnosis of model failures. The data generation pipeline combines automated scripts with human verification, ensuring high-quality, balanced samples.

Experimental results reveal that current open-source models like Qwen3-Omni perform poorly, with accuracy near 50% in most tasks, indicating a tendency to overestimate semantic alignment. In contrast, the commercial Gemini 2.5 Pro demonstrates significant improvements, achieving over 93% accuracy in global semantic tasks and outperforming others in temporal reasoning. These findings highlight a stark gap in cross-modal understanding, especially in complex semantic and temporal contexts.

The analysis suggests that the primary bottleneck lies in the models’ limited ability to fuse multimodal information effectively, despite strong single-modality perception. This work underscores the urgent need for developing models with enhanced cross-modal reasoning capabilities, particularly in speech-grounded video comprehension. Future research will focus on expanding datasets, incorporating multi-turn dialogues, and improving model robustness, ultimately aiming to create AI systems that reliably interpret complex multimedia content without hallucinating. This study provides a vital step toward safer and more trustworthy multimodal AI applications across industries such as video analysis, virtual assistants, and multimedia content creation.

Deep Analysis

Background

The evolution of multimodal AI has transitioned from early image captioning models like Show and Tell to sophisticated video understanding architectures such as VideoBERT and VideoLLaMA. These models leverage large-scale pretraining on datasets like MSR-VTT, YouCook2, and HowTo100M, achieving remarkable progress in tasks like captioning and question answering. However, a persistent challenge remains: models often generate outputs that are plausible but not grounded in the input data, a phenomenon known as hallucination. Existing benchmarks, such as VQA and MSRVTT-QA, mainly evaluate environmental sound detection or simple event recognition, neglecting the rich semantics and temporal complexities of human speech. As speech carries instructions, comments, and nuanced descriptions, understanding its alignment with visual content is crucial for real-world applications like video summarization, assistive technology, and autonomous systems. Despite advances, the gap in reliably integrating speech semantics with visual scenes persists, necessitating targeted evaluation tools.

Core Problem

Current multimodal models struggle with accurately aligning speech content with visual scenes, often hallucinating non-existent objects or events, especially when dealing with complex semantics and temporal relationships. This misalignment hampers their deployment in safety-critical applications, where incorrect interpretations can lead to failures or misinformation. Existing benchmarks lack the capacity to systematically evaluate these issues, particularly in the context of human speech, which involves dense, structured semantics and variable temporal references. Consequently, there is a pressing need for a dedicated, comprehensive evaluation framework to diagnose and address these limitations.

Innovation

The paper introduces SVHalluc, a novel benchmark that explicitly targets speech-vision hallucination across semantic and temporal dimensions. It innovates by:

  • �� Designing six multi-level tasks that probe model understanding from global scene matching to fine-grained object recognition and temporal reasoning.
  • �� Developing an automated data synthesis pipeline based on YouCook2, combined with human verification, to generate balanced positive and negative samples.
  • �� Incorporating diverse question formats—binary, multiple-choice, and event-specific—to comprehensively assess model capabilities.
  • �� Providing detailed diagnostic insights into failure modes, guiding future model improvements.

This approach advances beyond prior benchmarks by focusing on the rich semantics and complex temporal structures inherent in speech, addressing a critical gap in multimodal evaluation.

Methodology

  • �� Collect synchronized speech-video pairs from YouCook2, applying Whisper ASR for transcript generation.
  • �� Construct six tasks: GSA (global semantic alignment), FGSA (fine-grained object recognition), CMSB (semantic binding), TA (temporal alignment), TF (temporal forecasting), CMTB (temporal binding).
  • �� Generate positive samples using original aligned pairs; create negative samples by random pairing, object hiding, or event mismatching.
  • �� Use GPT-based prompts to extract objects and actions for negative sample creation, ensuring realistic combinations.
  • �� For temporal tasks, annotate event start/end times, define temporal relations, and generate questions accordingly.
  • �� Evaluate models with accuracy, precision, recall, and F1 metrics, applying human validation to ensure data quality.
  • �� Analyze failure modes, focusing on cross-modal understanding and temporal reasoning deficiencies.

Experiments

The evaluation involved six models: four open-source (Qwen3-Omni, Qwen2.5-Omni, VideoSALMONN 2, VideoLLaMA 2) and one commercial (Gemini 2.5 Pro). Zero-shot prompting was used across tasks, with questions designed to test semantic and temporal understanding. Metrics included accuracy for binary tasks and overall accuracy for multiple-choice questions. Results showed open-source models hovered around chance levels (~50%) in most tasks, especially in semantic alignment and temporal reasoning, indicating severe hallucination issues. Gemini 2.5 Pro significantly outperformed others, especially in global semantic and temporal tasks, validating the benchmark's effectiveness in exposing these deficiencies. Further analysis linked failures to poor cross-modal fusion despite strong unimodal perception.

Results

Open-source models' accuracy in semantic tasks was near 50%, indicating random guessing, while Gemini achieved over 93% in global semantic alignment. Fine-grained tasks revealed limited object recognition capabilities, with models often hallucinating non-existent entities. Temporal tasks showed most models failed to correctly identify event timing, with accuracy close to chance, confirming weak temporal reasoning. Gemini's superior performance highlights the importance of advanced cross-modal fusion mechanisms. These results underscore the critical need for improved multimodal understanding to prevent hallucinations in real-world applications.

Applications

The benchmark can guide development of safer, more reliable multimodal AI systems for video analysis, virtual assistants, and multimedia content moderation. It enables targeted diagnosis of hallucination modes, fostering models that accurately interpret complex speech-visual relationships. This has immediate implications for industries relying on accurate video content understanding, such as entertainment, security, and assistive technologies. Long-term, it promotes the creation of AI capable of nuanced, context-aware reasoning, reducing misinformation and enhancing human-AI collaboration in multimedia environments.

Limitations & Outlook

The current benchmark focuses on static question-answering tasks, lacking multi-turn dialogue or generative capabilities that are vital for interactive applications. Data diversity is limited to YouCook2, which may not reflect the full spectrum of real-world scenarios. The evaluation emphasizes accuracy metrics without assessing model robustness under adversarial or noisy conditions. Future work should incorporate dynamic, multi-turn tasks, larger datasets, and robustness testing to better simulate practical deployment environments.

Plain Language Accessible to non-experts

想象你在厨房里准备一顿饭,食材代表视频中的内容,调味料代表语音信息。模型就像厨师,要把调味料和食材配合得当,才能做出美味佳肴。有时候,厨师会搞错,把调味料放到错的食材上,或者把菜炒得不对时间。这就像模型产生的“幻觉”——它们会把语音和视频内容混淆,做出不存在的场景或事件。科学家们设计了一套测试,像是给厨师出题,看他们是否能正确配对调味料和食材,做出真正的菜。通过这个测试,我们可以知道模型在哪些方面容易出错,未来可以帮它们变得更聪明,做出更靠谱的“菜”。

ELI14 Explained like you're 14

你知道有时候看视频时,电脑会胡乱猜出一些根本没有的东西吗?比如视频里没有狗,但它说有狗在叫,或者说人正在做某件事,但其实根本没有。这叫做“幻觉”,就是模型自己想象出来的内容。最近的研究发现,这些模型在理解视频和语音时,经常会出现这种错误。科学家们设计了一些特别的测试,让模型回答关于视频和语音的问题,比如:视频里是不是有某个物体?讲述的事件是不是正在发生?结果发现,大部分开源模型表现很差,几乎和随机猜一样,而一些商业模型表现还不错。这说明,虽然模型在识别单一内容方面很厉害,但把语音和视频结合起来理解时,还差得远。未来,我们希望让模型变得更聪明,能更准确地理解视频内容,避免“幻觉”,让它们在实际应用中更可靠。

Abstract

Despite the success of audio-visual large-language models (LLMs), they can produce plausible but ungrounded outputs, termed hallucination. Existing benchmarks focus on environmental sounds (e.g., dog barking) to indicate event occurrence. In contrast, human speech carries fundamentally different, rich semantics and temporal structures, yet it remains unexplored whether current models can accurately align speech content with corresponding visual signals. In this work, we show that speech content can induce hallucinations in audio-visual LLMs. To systematically study this, we introduce SVHalluc, the first comprehensive benchmark for evaluating speech-vision hallucination in audio-visual LLMs. Our benchmark diagnoses speech-vision hallucinations from two critical and complementary aspects: semantic and temporal. Experimental results demonstrate that state-of-the-art open-source audio-visual LLMs struggle with aligning speech content with corresponding visual signals, with a near-random accuracy on multiple tasks. In contrast, Gemini 2.5 Pro significantly outperforms the open-source models. Our analysis suggests that their failures stem from limited ability in cross-modality understanding, despite strong performance in single-modality perception. Our work uncovers a new and fundamental limitation of current audio-visual LLMs and highlights the need for speech-grounded video comprehension. Project page: https://chenshuang-zhang.github.io/projects/svhalluc/.

eess.AS cs.AI cs.CV cs.LG cs.MM cs.SD