RefBench-PRO: Perceptual and Reasoning Oriented Benchmark for Referring Expression Comprehension
RefBench-PRO improves localization accuracy in referring expression comprehension using Ref-R1 with Dynamic IoU-based GRPO.
Key Findings
Methodology
RefBench-PRO benchmark decomposes referring expressions into perception and reasoning, further subdivided into six tasks: attribute, position, interaction, commonsense, relation, and reject. An automated data-generation pipeline produces diverse expressions. Ref-R1, incorporating Dynamic IoU-based GRPO, enhances localization accuracy under complex reasoning conditions.
Key Results
- RefBench-PRO poses greater challenges in perception and reasoning, with Ref-R1 significantly improving localization accuracy under complex conditions.
- On RefBench-PRO, models achieved 59.5% accuracy in visual-cue perception and 61.6% in relation and commonsense reasoning.
- Compared to existing benchmarks, RefBench-PRO excels in small target recognition and complex scenarios.
Significance
RefBench-PRO provides a comprehensive evaluation framework for multimodal large language models, revealing limitations across cognitive abilities. It enhances understanding of models' perception and reasoning capabilities and offers new insights for future model design.
Technical Contribution
RefBench-PRO enhances evaluation precision by refining task definitions and diversifying visual contexts. Ref-R1, with Dynamic IoU-based GRPO, introduces a new training strategy, improving model performance under complex reasoning conditions.
Novelty
RefBench-PRO is the first to decompose referring expression comprehension into perception and reasoning dimensions, generating diverse data through an automated pipeline for a more challenging benchmark.
Limitations
- Models show limited performance in the reject task, often producing localization hallucinations.
- Reasoning capabilities in complex scenarios need further improvement.
Future Work
Future work could optimize Ref-R1's training strategy to enhance model performance across tasks and explore more application scenarios.
AI Executive Summary
RefBench-PRO is a benchmark focused on perception and reasoning in referring expression comprehension, addressing the lack of interpretable scoring mechanisms in existing benchmarks. By decomposing referring expressions into perception and reasoning, further subdivided into six tasks, RefBench-PRO offers a comprehensive evaluation framework.
To generate diverse referring expressions, the research team developed an automated data-generation pipeline capable of producing diverse data across six sub-dimensions. Additionally, the proposed Ref-R1 learning scheme, incorporating Dynamic IoU-based GRPO, significantly improves localization accuracy under complex reasoning conditions.
The experimental results demonstrate that RefBench-PRO reveals the limitations of multimodal large language models in referring expression comprehension, offering new insights for future model design. Despite limited performance in the reject task, RefBench-PRO is significant in enhancing models' perception and reasoning capabilities.
Deep Analysis
Background
Referring Expression Comprehension (REC) is a vision-language task aimed at localizing specific image regions based on textual descriptions. Existing REC benchmarks primarily evaluate perceptual capabilities but lack interpretable scoring mechanisms, failing to reveal the grounding capabilities of multimodal large language models across different cognitive abilities.
Core Problem
Existing benchmarks inadequately evaluate the perception and reasoning capabilities of multimodal large language models, particularly in visually complex scenes and compositional reasoning. A comprehensive benchmark is needed to evaluate model capabilities.
Innovation
RefBench-PRO decomposes referring expressions into perception and reasoning, subdivided into six tasks, providing a comprehensive evaluation framework. The automated data-generation pipeline and Ref-R1 learning scheme are core innovations.
Methodology
- �� Decompose referring expressions into perception and reasoning dimensions.
- �� Develop an automated data-generation pipeline for diverse expressions.
- �� Propose Ref-R1 learning scheme with Dynamic IoU-based GRPO to enhance localization accuracy.
Experiments
Experiments use the RefBench-PRO benchmark to evaluate multimodal large language models across six tasks. Dynamic IoU-based GRPO is used for training, comparing models' perception and reasoning capabilities.
Results
Experiments show that RefBench-PRO poses greater challenges in perception and reasoning. Ref-R1 significantly improves localization accuracy under complex reasoning conditions.
Applications
RefBench-PRO can be used to evaluate multimodal large language models in referring expression comprehension, applicable in scenarios requiring precise localization and complex reasoning, such as autonomous driving and intelligent surveillance.
Limitations & Outlook
While RefBench-PRO is significant in enhancing models' perception and reasoning capabilities, models show limited performance in the reject task, often producing localization hallucinations.
Plain Language Accessible to non-experts
Imagine you're in a large supermarket, needing to find a specific item. RefBench-PRO is like a super-smart shopping assistant that not only finds the item based on your description but also accurately locates it in a complex shelf layout. It analyzes the item's attributes, position, and relationship with other items to help you quickly find your target. Even if the item is out of stock, it can tell you it's not on the shelf.
ELI14 Explained like you're 14
Imagine you're playing a treasure hunt game, and RefBench-PRO is your super assistant. It helps you find the treasure based on clues, even if the clues are complex. It analyzes each clue's details, like the treasure's color, position, and relationship with other objects. Even if the treasure isn't on the map, it can tell you not to waste time. Isn't that cool?
Glossary
Referring Expression Comprehension
A vision-language task aimed at localizing specific image regions based on textual descriptions.
Used in the paper to evaluate the perception and reasoning capabilities of multimodal large language models.
Multi-modal Large Language Model
Large models that combine visual and language information to handle complex vision-language tasks.
Used for referring expression comprehension tasks.
Dynamic IoU-based GRPO
A reinforcement learning strategy that combines Dynamic IoU to enhance localization accuracy.
Used in the Ref-R1 learning scheme to improve localization accuracy under complex reasoning conditions.
Perception
The ability of a model to recognize and understand visual cues.
Used as one of the evaluation dimensions in RefBench-PRO.
Reasoning
The ability of a model to analyze and integrate visual and textual information to draw conclusions.
Used as one of the evaluation dimensions in RefBench-PRO.
Open Questions Unanswered questions from this research
- 1 How to improve model performance in the reject task to avoid localization hallucinations?
- 2 How to further enhance reasoning capabilities in complex scenarios?
Applications
Immediate Applications
Autonomous Driving
RefBench-PRO can be used to evaluate autonomous driving systems' perception and decision-making capabilities in complex scenarios.
Long-term Vision
Intelligent Surveillance
RefBench-PRO can be used to develop smarter surveillance systems, enhancing anomaly detection and response capabilities.
Abstract
Referring Expression Comprehension (REC) is a vision-language task that localizes a specific image region based on a textual description. Existing REC benchmarks primarily evaluate perceptual capabilities and lack interpretable scoring mechanisms, which cannot reveal the grounding capability of Multi-modal Large Language Model (MLLM) across different cognitive abilities. To address this limitation, we introduce RefBench-PRO, a comprehensive REC benchmark, which decomposes referring expressions into two core dimensions, i.e., perception and reasoning, and further subdivides them into six progressively challenging tasks, such as attribute, position, interaction, commonsense, relation and reject. We also develop a fully automated data-generation pipeline that produces diverse referring expressions across these six sub-dimensions. Furthermore, We propose Ref-R1, an RL-based learning scheme, which incorporates Dynamic IoU-based GRPO to improve localization accuracy under increasingly complex reasoning conditions, establishing a stronger baseline for REC. Extensive experiments demonstrate that our RefBench-PRO enables interpretable evaluation of MLLM on referring expression comprehension, presenting greater challenges in both perception and reasoning.