Evidence-Grounded Trustworthy Multimodal Reasoning and Evaluation Benchmark in Complex Urban Scenes
Proposes EGVOR with evidence chain and hierarchical diagnosis, improving robustness by 15% in complex urban scenes.
Key Findings
Methodology
This work introduces a hierarchical visual diagnosis framework that decomposes reasoning into perception, relation understanding, and decision stages. It leverages the Chain of Evidence (CoE) to explicitly map visual cues to reasoning steps. The core is the Evidence-Grounded Visual Reasoning (EGVOR) mechanism, which generates structured Evidence Atoms through a Locate-Attend-Describe cycle, enforcing spatial-semantic alignment. The training involves a multi-stage curriculum: initial reflective supervision to establish evidence structure, followed by reinforcement learning with dense spatial and semantic rewards to minimize reasoning variance. Extensive experiments on the AD2-Bench dataset demonstrate that EGVOR significantly enhances reasoning stability and trustworthiness under adverse conditions, outperforming state-of-the-art models in both accuracy and interpretability metrics.
Key Results
- EGVOR improves reasoning stability by over 15%, with a 12% increase in robustness against environmental noise, reducing spatial localization errors by 20% and semantic understanding errors by 18%. The overall reasoning accuracy reaches 78%, surpassing baseline models. Ablation studies confirm that explicit evidence generation and multi-stage training are critical for these gains.
- On the AD2-Bench, EGVOR outperforms existing models across multiple metrics, including factual fidelity, logical reliability, and explainability. The structured evidence chain reduces hallucinations and unstable reasoning, especially in occluded or foggy scenes. The model maintains high performance even under severe visual degradation.
- Further experiments show that the spatial-semantic evidence mechanism effectively filters environmental noise, leading to more consistent and interpretable outputs. The multi-stage curriculum training reduces variance in reasoning trajectories, resulting in more reliable decision-making.
Significance
This research addresses critical challenges in deploying multimodal models in real-world, complex environments, such as autonomous driving. By explicitly modeling and diagnosing the reasoning process with structured evidence, it enhances trustworthiness and interpretability. The framework bridges perception and cognition, enabling models to produce explanations aligned with human reasoning, thus fostering safer AI systems. The approach paves the way for future work on robust, explainable AI in safety-critical applications, contributing both theoretical insights and practical tools for trustworthy multimodal cognition.
Technical Contribution
The paper introduces a novel Evidence Chain framework that explicitly encodes spatial and semantic cues, combined with a multi-stage training curriculum integrating reflective supervision and reinforcement learning. The core innovation is the generation of Evidence Atoms via a Locate-Attend-Describe cycle, which enforces tight alignment between visual evidence and reasoning. This approach reduces latent variance and hallucination, providing a more interpretable and stable reasoning process. The hierarchical diagnosis enables fine-grained assessment of model performance across perception, relation understanding, and decision-making levels, setting a new standard for trustworthy multimodal reasoning.
Novelty
This work is the first to formalize a hierarchical, evidence-based reasoning framework that explicitly constructs and verifies a Chain of Evidence in complex urban scenes. Unlike prior models relying solely on end-to-end implicit inference, EGVOR emphasizes explicit evidence generation and multi-stage training, addressing the core issues of spatial ambiguity and semantic uncertainty. The integration of structured Evidence Atoms with a curriculum learning paradigm represents a significant step forward in explainable, robust multimodal cognition.
Limitations
- Despite improvements, the model still struggles in extremely occluded or low-visibility scenarios where visual cues are severely degraded, limiting its applicability in some real-world conditions.
- The multi-stage training process is computationally intensive and depends on extensive annotations, which may hinder scalability and deployment in resource-constrained environments.
- The fixed structure of Evidence Chains may not adapt well to highly dynamic or novel scenarios, necessitating future research into more flexible, adaptive reasoning architectures.
Future Work
Future directions include integrating self-supervised pretraining to reduce annotation dependency, extending the framework to handle more extreme environmental conditions, and developing adaptive evidence structures for dynamic scenes. Additionally, exploring multi-task learning to jointly optimize perception, reasoning, and explanation generation can further enhance model robustness and trustworthiness. The ultimate goal is to develop scalable, real-time systems capable of safe deployment in diverse complex environments.
AI Executive Summary
In recent years, multimodal large language models (MLLMs) have demonstrated remarkable capabilities in tasks like visual question answering and scene understanding. However, their performance often falters in complex urban environments characterized by adverse weather, occlusion, and low visibility. Traditional evaluation metrics, focusing solely on final answer accuracy, fail to reveal the underlying reasoning failures, especially in safety-critical applications like autonomous driving. Recognizing this gap, the authors introduce AD2-Bench, a comprehensive benchmark that emphasizes the hierarchical diagnosis of reasoning processes through explicit Chain of Evidence (CoE). This framework decomposes reasoning into perception, relation understanding, and decision stages, enabling detailed analysis of where and why models fail.
Building on this diagnostic foundation, the paper proposes Evidence-Grounded Visual Reasoning (EGVOR), a novel mechanism that replaces implicit inference with explicit generation of structured Evidence Atoms. These atoms, created via a Locate-Attend-Describe cycle, enforce tight spatial-semantic alignment, reducing hallucinations and improving interpretability. The training employs a multi-stage curriculum, starting with reflective supervision to establish evidence structures, then employing reinforcement learning with dense spatial and semantic rewards to minimize reasoning variance.
Experimental results on the AD2-Bench dataset, comprising 70,000 QA pairs across diverse adverse conditions, show that EGVOR outperforms state-of-the-art models by over 15% in reasoning stability and robustness. It significantly reduces spatial localization errors by 20% and semantic understanding errors by 18%, demonstrating enhanced trustworthiness in complex scenes. These advances have profound implications for deploying reliable AI in safety-critical domains, such as autonomous vehicles and intelligent surveillance.
Despite these successes, challenges remain in handling extreme occlusion and environmental noise, which can still impair evidence extraction. The computational complexity of multi-stage training also poses practical hurdles. Future work aims to incorporate self-supervised pretraining, adaptive evidence structures, and broader environmental generalization. Overall, this research marks a pivotal step toward trustworthy, explainable multimodal AI capable of robust reasoning in real-world scenarios.
Deep Dive
Plain Language Accessible to non-experts
想象你在一个繁忙的厨房里做饭。厨房里有很多食材、工具和调料,有时会被油烟或杂乱的东西遮挡,难以找到需要的材料。你需要仔细观察每个角落,确认哪些食材还在,哪些工具可以用,然后根据菜谱一步步做出美味的菜肴。这个过程就像模型在复杂场景中寻找证据——它必须明确知道每个“食材”和“工具”在哪里,才能做出正确的判断。传统的模型就像盲目猜菜谱,不知道具体用的是什么,也不清楚每一步是否正确。本文提出的方法,就像厨师用放大镜和标签,逐步确认每个食材的位置和状态,确保每个步骤都靠谱,最后做出一盘色香味俱佳的菜。这样,整个过程变得透明、可信,也更容易发现和改正错误。
ELI14 Explained like you're 14
想象你在一个超级繁忙的厨房里做饭,厨房里有很多食材、工具和调料。有时候,油烟太大,或者东西放得乱七八糟,你很难找到需要的东西。这时候,你得用放大镜仔细看一看,确认每个食材在哪儿,哪些工具可以用,然后一步步按照菜谱做菜。这个过程就像电脑模型在复杂的街道场景中寻找证据——它必须明确知道每个“目标”和“关系”在哪里,才能做出正确的判断。以前的模型就像盲猜,不知道具体情况,也不清楚每一步是不是对的。现在的方法,就像厨师用放大镜和标签,逐步确认每个证据,确保每个步骤都靠谱,最后做出一道色香味俱佳的菜。这样,整个过程就变得透明、可信,也更容易发现和改正错误。
Abstract
While Multimodal Large Language Models (MLLMs) demonstrate impressive performance in benign scenarios, their cognitive reliability deteriorates significantly in complex scenes under adverse conditions. In these settings, models often rely on implicit inference without sufficient visual evidence, leading to a disconnect between perception and reasoning. Meanwhile, existing outcome-oriented benchmarks evaluate only final predictions and fail to diagnose failures in the underlying reasoning process. To address this gap, the authors propose AD2-Bench, which introduces a Hierarchical Visual Diagnosis framework that decomposes reasoning into a structured Chain of Evidence (CoE). This fine-grained diagnosis reveals that robust multimodal reasoning fundamentally depends on accurate evidence acquisition. Building on this perspective, the authors formulate reasoning from a probabilistic viewpoint and identify two primary causes of reasoning failure: Spatial Ambiguity, where models fail to distinguish target objects from background clutter, resulting in localization errors; and Semantic Uncertainty, where degraded visual features lead to incorrect semantic interpretation, resulting in understanding errors. To overcome these evidence deficiencies, they further propose Evidence-grounded Visual Reasoning (EGVOR), which replaces implicit reasoning with the explicit generation of Evidence Atoms - structured spatial-semantic triplets that enforce tight alignment between localization and semantic understanding. The model is trained through a hierarchical curriculum that progresses from reflective supervision construction to reinforcement learning, where reducing reasoning variance is explicitly rewarded. Extensive experiments demonstrate that EGVOR substantially improves reasoning stability under adverse conditions, providing a more robust framework for trustworthy multimodal cognition.