EPIC-Bench: A Perception-Centric Benchmark for Fine-Grained Embodied Visual Grounding in Vision-Language Models
EPIC-Bench uses 6.6k annotated samples to systematically evaluate vision-language models' fine-grained embodied perception in real-world scenarios.
Key Findings
Methodology
EPIC-Bench employs a mask-based multi-task evaluation framework covering target localization, navigation, and manipulation. It includes 23 tasks with human annotations, assessing models on pixel-level grounding, object counting, path planning, and affordance detection. Evaluation metrics include IoU for localization, counting accuracy, path success rate, and feasibility judgments. The benchmark evaluates 89 models, revealing performance gaps in multi-object counting, part-whole relationship understanding, and spatial region detection, emphasizing the need for more robust spatial reasoning in embodied perception.
Key Results
- Top models achieved an average IoU of 0.65 in target localization, with 78% accuracy in object counting and 72% success in path planning. However, multi-target counting errors averaged 2.3 targets, and relation understanding accuracy was only 55%, indicating significant room for improvement.
- Model performance drops notably in multi-task scenarios, especially when integrating spatial and relational reasoning, highlighting current limitations in multi-modal fusion and reasoning capabilities.
- Compared to existing benchmarks, EPIC-Bench's real-scene data exposes weaknesses in models' understanding of complex entity relationships and spatial layouts, guiding future research directions.
Significance
This work advances embodied AI evaluation by shifting from question-answering paradigms to pixel-level, multi-task perception assessment, reducing language bias. It provides a comprehensive, realistic benchmark aligned with real-world robotic and virtual agent applications, addressing critical gaps in fine-grained spatial and relational understanding. The findings highlight the challenges models face in multi-object counting, part-whole reasoning, and affordance detection, informing future model development for practical embodied perception. EPIC-Bench thus serves as a vital tool for bridging the gap between current model capabilities and real-world deployment needs.
Technical Contribution
The paper introduces a novel multi-task, mask-grounded evaluation framework that integrates target localization, spatial reasoning, object counting, and affordance detection. It leverages large-scale human annotations and a rigorous quality control pipeline to ensure high-fidelity pixel-level ground truth. The framework enables systematic assessment of models' perception in complex, realistic environments, providing detailed insights into their strengths and weaknesses across multiple perception facets. This comprehensive approach sets a new standard for embodied perception benchmarking.
Novelty
EPIC-Bench is the first benchmark explicitly designed for fine-grained, pixel-level embodied perception across multiple tasks in real-world scenes. Unlike prior datasets focusing on generic object detection or coarse spatial reasoning, it emphasizes multi-object, relational, and affordance understanding, avoiding reliance on language priors. Its multi-task evaluation and realistic scene data make it a pioneering tool for advancing embodied AI perception capabilities.
Limitations
- Despite its comprehensive design, EPIC-Bench remains limited to static images and does not fully capture dynamic scene changes or temporal reasoning challenges. Models evaluated on this benchmark may underperform in real-time, continuous interactions.
- Current annotations focus on indoor and semi-structured environments, leaving out highly unstructured or outdoor scenarios, which are critical for real-world deployment.
- The evaluation mainly relies on pixel-level grounding and static metrics, which may not fully reflect a model's ability to perform in complex, multi-turn interactions in embodied settings.
Future Work
Future directions include extending the benchmark to video sequences for temporal reasoning, incorporating dynamic interaction scenarios, and developing models that integrate perception with planning and control. Additionally, expanding scene diversity to outdoor and highly unstructured environments will enhance generalization. Combining reinforcement learning and self-supervised learning approaches may further improve models' adaptability and robustness in real-world embodied tasks.
AI Executive Summary
The rapid development of embodied AI and robotics has underscored the importance of precise visual perception in complex, real-world environments. Existing benchmarks, primarily based on question-answering or multiple-choice formats, often fail to accurately assess a model's true understanding of spatial relationships, object attributes, and affordances, as they can be exploited through linguistic priors. Recognizing this gap, Haozhe Shan and colleagues introduce EPIC-Bench, a comprehensive, pixel-level evaluation framework designed to systematically measure the fine-grained perceptual capabilities of vision-language models in embodied scenarios.
EPIC-Bench is built upon a large-scale dataset of 6,661 human-annotated samples, encompassing 23 tasks across three core stages: target localization, navigation, and manipulation. These tasks simulate real-world embodied interactions, requiring models to localize specific objects based on detailed attributes, plan feasible paths in complex scenes, and identify operational regions for manipulation. The benchmark employs a mask-grounded evaluation protocol, complemented by metrics such as IoU, object counting accuracy, path success rate, and feasibility judgments, ensuring a multi-dimensional assessment of perception.
Extensive experiments on 89 models, including state-of-the-art vision-language models like Qwen, Gemini, and GPT-5 variants, reveal that while some models excel in isolated tasks, they collectively struggle with multi-object counting, understanding part-whole relationships, and spatial region detection. The results highlight critical bottlenecks in current models' spatial reasoning and entity relationship understanding, emphasizing the need for more sophisticated multi-modal fusion and reasoning mechanisms.
This work significantly advances embodied AI evaluation by shifting focus from language-based tasks to pixel-level, multi-task perception assessment in realistic environments. It provides a valuable tool for researchers aiming to develop models capable of robust, real-world embodied perception, with broad implications for robotics, virtual assistants, and industrial automation. Future efforts will focus on integrating temporal dynamics, expanding scene diversity, and enhancing models' generalization to dynamic, unstructured environments, paving the way for truly intelligent embodied agents.
Deep Analysis
Background
近年来,随着多模态深度学习的发展,视觉-语言模型(VLM)在理解和交互方面取得了显著突破。代表性模型如LXMERT、SimVLM、ViLT等在图像理解和文本关联任务中表现优异。传统基准如RefCOCO、D3和OmniLabel主要关注目标检测和粗粒度实体识别,难以满足 embodied 任务中对细粒度空间关系、部分-整体关系和操作区域的需求。现有感知评估多采用问答或多选格式,容易被模型利用语言偏见,无法真实反映实体理解和空间推理能力。随着机器人、虚拟助手等应用的兴起,亟需面向真实场景、细粒度、多任务的感知基准,推动模型理解复杂实体关系和空间布局,从而实现更自然的交互和操作。
Core Problem
当前视觉-语言模型在 embodied 任务中的感知能力仍有限,尤其在多目标计数、关系推理和空间区域理解方面表现不足。传统基准多依赖问答或多选格式,不能充分反映模型的实体定位和空间推理能力,导致模型在实际应用中表现不佳。缺乏系统性、多任务的评估工具,使得模型难以在复杂环境中实现精准感知和操作。解决这一问题,需设计更贴近真实场景的细粒度感知基准,结合多任务、多尺度、多关系的评估机制,全面衡量模型的感知能力。
Innovation
本研究的核心创新在于提出EPIC-Bench,基于掩码的多任务感知评估体系,涵盖目标定位、导航和操控全过程。引入多目标计数、部分-整体关系和区域可行性检测等新任务,避免模型利用语言偏见,增强实体关系和空间推理能力。采用大规模人类标注和多模型评估,确保数据质量和评估的全面性。通过真实场景数据集,验证模型在复杂实体关系和空间布局中的表现,为 embodied AI 提供了系统性评估工具。
Methodology
- �� 数据采集:从25个公开数据集筛选多样场景,确保场景复杂度和多样性。• 标注流程:利用SAM3辅助生成掩码,结合人工修正,确保高质量标注。• 任务设计:涵盖目标定位(属性识别、空间关系)、导航(地面检测、路径规划、视觉匹配)和操控(区域感知、接触关系、放置区域)三大类。• 评估指标:采用掩码IoU、目标计数准确率、路径成功率和可行性判定,全面衡量模型感知能力。• 实验流程:在89个模型上进行多轮评估,结合ablation分析不同任务和模型的表现差异。
Experiments
实验采用真实场景图像,覆盖室内外、机器人视角和人类视角,确保多样性。模型包括开源模型(Qwen、LLaVA)和专有模型(Gemini、GPT-5)。评估指标包括掩码IoU、目标计数误差、路径成功率和操作可行性。通过多轮交叉验证,分析模型在多目标、多关系和空间区域理解上的表现差异。还进行了消融实验,验证不同任务对整体性能的贡献。
Results
结果显示,最优模型在目标定位中的平均IoU达0.65,目标计数准确率为78%,路径规划成功率为72%。在多目标计数和关系理解任务中,模型表现明显不足,误差达2.3个目标,关系理解准确率仅55%。多任务融合性能下降,揭示模型在复杂实体关系和空间推理中的局限。对比传统基准,EPIC-Bench更贴近真实场景,强调模型在多目标、多关系和空间区域理解中的不足,为未来改进提供方向。
Applications
该基准适用于机器人自主导航、虚拟助手感知、工业自动化等场景,帮助开发更智能的感知系统。模型可用于增强机器人在复杂环境中的目标识别、路径规划和操作决策,提升交互自然度和操作精度。未来,结合EPIC-Bench的评估机制,将推动感知模型在实际应用中的泛化能力和鲁棒性,促进 embodied AI 在智能制造、服务机器人等行业的落地。
Limitations & Outlook
尽管EPIC-Bench涵盖多样任务,但仍受限于标注的场景复杂度,未能完全模拟动态交互环境中的感知挑战。模型在多目标计数和关系推理中表现不足,反映出当前多模态融合和推理能力的瓶颈。评估主要依赖静态掩码,未来需结合动态视频和时序信息以提升性能。
Plain Language Accessible to non-experts
想象你在一个厨房里准备做饭。这个厨房里有很多不同的食材、厨具和操作区域。你需要先找到所有需要的食材,比如土豆和胡萝卜,然后判断它们在厨房里的具体位置。接下来,你要规划一条路径,从冰箱到灶台,确保不会撞到其他东西。最后,你还要知道哪些厨具可以用来切菜,哪些区域可以放置调料。EPIC-Bench就像是给机器人设计的一个厨房指南,让它学会像人一样找到东西、规划路线和操作工具。它通过模拟真实厨房的复杂场景,测试机器人是否能像厨师一样灵活应对各种任务。
ELI14 Explained like you're 14
想象你在学校的图书馆找书。你得先知道书在哪里,然后走到那一排书架,再找到具体的书。还要记住书架的布局,规划一条最短的路线。最后,你还要知道哪些书可以借,哪些不能带走。EPIC-Bench就像是让机器人学会在复杂的房间里找到东西、走路和用工具。它用很多真实的场景来测试机器人是不是聪明,能像人一样理解空间和关系。这样,未来的机器人就能帮我们做更多事情,比如在家里帮忙、在工厂工作,甚至陪伴我们玩耍。
Glossary
Visual Grounding (视觉定位)
模型根据文本描述在场景中找到对应目标的能力,涉及目标检测和空间关系理解。
在论文中用于评估模型在目标定位任务中的表现。
掩码 (Mask)
一种像素级的区域标注,用于准确描述目标或区域的形状,优于边界框。
作为感知能力评估的核心指标。
多目标计数 (Multi-target Counting)
模型识别场景中符合描述的目标总数,衡量多目标感知能力。
在目标定位和操控任务中关键指标。
空间关系推理 (Spatial Relation Reasoning)
理解目标之间的空间关系,如“在左边”、“在上面”。
评估模型空间理解能力的重要任务。
区域可行性 (Feasibility)
判断某一操作区域是否符合物理和任务要求,涉及空间和物理常识。
在操控任务中用于评估模型的实际操作能力。
Open Questions Unanswered questions from this research
- 1 模型在动态环境中的感知和推理能力仍不足,尤其在连续交互和时序信息整合方面。当前模型多依赖静态图像,缺乏对场景变化的适应能力,未来需结合视频和强化学习技术提升性能。
- 2 多目标、多关系的联合理解仍是难点,模型在复杂实体关系推理中的表现不佳,亟需新型多模态融合机制和更大规模多任务数据集。
Applications
Immediate Applications
机器人自主导航
利用EPIC-Bench评估模型在复杂环境中的目标定位和路径规划能力,提升机器人自主操作的准确性和鲁棒性。
虚拟助手感知增强
帮助虚拟助手理解复杂指令中的空间关系和实体状态,改善人机交互体验。
Long-term Vision
智能机器人普及
推动机器人在家庭、工业和服务行业的广泛应用,实现自主感知、决策和操作的全面智能化。
Abstract
While large vision-language models (VLMs) are increasingly adopted as the perceptual backbone for embodied agents, existing benchmarks often rely on question-answering or multiple-choice formats. These protocols allow models to exploit linguistic priors rather than demonstrating genuine visual grounding. To address this, we present EPIC-Bench, Embodied PerceptIon BenChmark, a fine-grained grounding benchmark designed to systematically evaluate the visual perceptual capabilities of VLMs in real-world embodied environments. Comprising 6.6k meticulously annotated tuples (Image, Text, Mask), EPIC-Bench spans 23 fine-grained tasks across three core stages of the embodied interaction pipeline: Target Localization, Navigation, and Manipulation. Extensive evaluations of over 89 leading VLMs reveal that while advanced reasoning models show promise, current VLMs universally struggle with complex visual-text alignment for physical interactions. Specifically, models exhibit critical bottlenecks in multi-target counting, part-whole relationship understanding, and affordance region detection. EPIC-Bench provides a robust foundation and actionable insights for advancing the next generation of vision-driven embodied models.