VANTAGE-Bench: Evaluating the Infrastructure AI Gap in Vision-Language Models

TL;DR

VANTAGE-Bench evaluates the Infrastructure AI gap in VLMs, revealing deficiencies in event verification and temporal localization tasks.

cs.CV 🔴 Advanced 2026-09-09 4 views
Zaid Pervaiz Bhat Nimra Nayyar Arihant Jain Lap Fung Chan John Suchanek Yu Wang Varun Praveen Tomasz Kornuta Vidya Nariyambut Murali
Vision-Language Models Infrastructure AI Benchmark Event Verification Temporal Localization

Key Findings

Methodology

VANTAGE-Bench is a multi-task benchmark specifically designed for Infrastructure AI, covering logistics, transportation, and smart spaces. It evaluates models' semantic, spatial, temporal, and spatio-temporal understanding through eight task formulations, moving beyond traditional multiple-choice formats, and introduces a single-pass trajectory-generation protocol for Single Object Tracking.

Key Results

  • In zero-shot evaluation of 17 models, event verification, referring expressions, and temporal localization tasks perform 9 to 24 points lower than consumer video benchmarks, while video QA is only 5.3 points lower.
  • Temporal localization tasks do not exceed 55.7 mIoU, and dense video captioning does not exceed 37.3 SODA_c.
  • Frontier models are within 5 points of specialist trackers in short-term tracking but diverge significantly in long-term tracking.

Significance

This study reveals the inadequacies of VLMs in fixed-camera environments, emphasizing the importance of evaluating models in operational settings. By providing a multi-task benchmark, VANTAGE-Bench offers new directions for research and development in Infrastructure AI, impacting fields like safety monitoring and operational logging.

Technical Contribution

VANTAGE-Bench fills the gap in existing benchmarks for Infrastructure AI with a multi-task evaluation framework. It provides an open evaluation ecosystem and a live public leaderboard, supporting comprehensive assessment of models in real-world operational environments.

Novelty

VANTAGE-Bench is the first multi-task benchmark specifically curated for Infrastructure AI, introducing the first evaluation of Single Object Tracking on fixed-camera video, breaking the traditional multiple-choice limitation.

Limitations

  • Models perform weakly in temporal localization and dense video captioning tasks, not exceeding 55.7 mIoU and 37.3 SODA_c.
  • Models diverge significantly from specialist trackers in long-term tracking tasks.

Future Work

Future research can focus on improving model performance in temporal localization and dense video captioning tasks, as well as enhancing robustness in long-term tracking tasks.

AI Executive Summary

As Vision-Language Models (VLMs) rapidly advance, existing research has focused primarily on action-oriented AI based on consumer videos, overlooking the critical domain of Infrastructure AI. Infrastructure AI relies on fixed cameras for safety monitoring and operational logging, with dense data and fixed perspectives that current internet video data cannot adequately evaluate.

VANTAGE-Bench is a benchmark specifically designed for Infrastructure AI, covering logistics, transportation, and smart spaces. It evaluates models' semantic, spatial, temporal, and spatio-temporal understanding through eight task formulations, moving beyond traditional multiple-choice formats, and introduces a single-pass trajectory-generation protocol for Single Object Tracking.

In zero-shot evaluation of 17 models, VANTAGE-Bench reveals deficiencies in event verification, referring expressions, and temporal localization tasks. While models perform closely to existing benchmarks in video QA, they remain weak in temporal localization and dense video captioning tasks. This study offers new directions for the development of Infrastructure AI.

Deep Analysis

Background

Vision-Language Models (VLMs) have made significant progress in recent years, primarily applied to action-oriented AI based on consumer videos. However, Infrastructure AI, which relies on fixed cameras for safety monitoring and operational logging, remains an important domain. Existing internet video data cannot adequately evaluate its performance, leading to a gap in model performance in real-world operational environments.

Core Problem

The core problem is the inadequate performance of existing Vision-Language Models in Infrastructure AI, particularly in event verification and temporal localization tasks. These tasks require high levels of semantic, spatial, and temporal understanding, posing higher demands on the evaluation framework.

Innovation

VANTAGE-Bench fills the gap in existing benchmarks for Infrastructure AI with a multi-task evaluation framework. It introduces the first evaluation of Single Object Tracking on fixed-camera video, breaking the traditional multiple-choice limitation, and offers a new method for assessing model performance in real-world operational environments.

Methodology

  • �� VANTAGE-Bench covers logistics, transportation, and smart spaces.
  • �� Provides eight task formulations, including semantic, spatial, temporal, and spatio-temporal understanding.
  • �� Introduces a single-pass trajectory-generation protocol for Single Object Tracking.
  • �� Dataset includes 3,346 media assets, covering video tasks, image grounding, and detection boxes.

Experiments

The experimental design includes zero-shot evaluation of 17 models, covering open-weight and proprietary models. Inference and metric computation were conducted using an extended VLMEvalKit harness, focusing on model performance in event verification, referring expressions, and temporal localization tasks.

Results

In zero-shot evaluation of 17 models, event verification, referring expressions, and temporal localization tasks perform 9 to 24 points lower than consumer video benchmarks, while video QA is only 5.3 points lower. Temporal localization tasks do not exceed 55.7 mIoU, and dense video captioning does not exceed 37.3 SODA_c.

Applications

VANTAGE-Bench's application scenarios include safety monitoring and operational logging in Infrastructure AI. By providing a multi-task evaluation framework, VANTAGE-Bench offers new directions for research and development in Infrastructure AI.

Limitations & Outlook

Models perform weakly in temporal localization and dense video captioning tasks, not exceeding 55.7 mIoU and 37.3 SODA_c. Models diverge significantly from specialist trackers in long-term tracking tasks.

Plain Language Accessible to non-experts

Imagine a large warehouse with many cameras monitoring every corner. These cameras are like the eyes of the warehouse, responsible for observing and recording all activities. VANTAGE-Bench is like a tool that tests these eyes to see if they can accurately identify and describe what happens. For example, when a worker moves goods, the cameras need to know where he started, where he went, and where he placed the goods. VANTAGE-Bench evaluates these cameras' performance through a series of complex tasks, ensuring they work effectively in real operations.

ELI14 Explained like you're 14

Imagine you're playing a super complex game with lots of tasks to complete, like finding out who's stealing or when an accident happened. VANTAGE-Bench is like the game's referee, helping us judge if these tasks are done well. It asks questions like 'Did this person use a key card to enter?' or 'When did the collision happen?' and checks if the model can give the right answers. This tool is super cool because it helps ensure that surveillance cameras work effectively in real life, keeping us safe.

Glossary

Vision-Language Model (VLM)

A Vision-Language Model is an AI model that combines image and text information for understanding and generation.

VANTAGE-Bench evaluates VLM performance in Infrastructure AI.

Infrastructure AI

Infrastructure AI relies on fixed cameras for safety monitoring and operational logging, providing open-loop insights.

VANTAGE-Bench is specifically designed to evaluate Infrastructure AI.

Event Verification

Event verification tasks require models to verify operational hypotheses against visual evidence.

VANTAGE-Bench evaluates model performance in event verification tasks.

Temporal Localization

Temporal localization tasks require models to accurately predict the start and end timestamps of actions.

VANTAGE-Bench evaluates model performance in temporal localization tasks.

Single Object Tracking

Single Object Tracking tasks require models to predict a target's coordinate trajectory across a full sequence.

VANTAGE-Bench introduces the first evaluation of Single Object Tracking on fixed-camera video.

Open Questions Unanswered questions from this research

  • 1 How to improve model robustness in long-term tracking tasks? Current models perform poorly in long-term tracking tasks, requiring exploration of more effective tracking algorithms.
  • 2 How to enhance model performance in temporal localization tasks? Current models do not exceed 55.7 mIoU in temporal localization tasks, necessitating the development of more precise temporal prediction algorithms.

Applications

Immediate Applications

Safety Monitoring

VANTAGE-Bench can be used to evaluate surveillance camera performance in safety monitoring, ensuring accurate identification and recording of safety events.

Operational Logging

By evaluating model performance in operational logging tasks, VANTAGE-Bench can help optimize operational efficiency in warehouses and transportation systems.

Long-term Vision

Smart Cities

VANTAGE-Bench can be used to evaluate Infrastructure AI systems in smart cities, promoting intelligent and automated urban management.

Abstract

As Vision-Language Models (VLMs) advance toward physical deployment, the focus has remained on action-oriented Embodied AI evaluated on subject-centric consumer video. This overlooks a pervasive class of Physical AI: Infrastructure AI, which relies on fixed cameras for open-loop insights like safety monitoring and operational logging. We introduce VANTAGE-Bench, a benchmark measuring this "Infrastructure AI Gap." It spans three operational domains (Logistics, Transportation, and Smart Spaces), unifies image and video evaluation across semantic, spatial, temporal, and spatio-temporal capabilities, and moves beyond multiple-choice to eight task formulations including dense captioning and spatio-temporal grounding. It adds a single-pass trajectory protocol for Single Object Tracking and, to our knowledge, the first such evaluation on fixed-camera infrastructure video, scored against specialist trackers. Annotation spans three regimes over 3,346 media assets: 3,342 video-task annotations, 4,281 image-grounding annotations, and 27,404 detection boxes. Evaluating 17 models zero-shot, we find the shortfall relative to consumer-centric benchmarks is concentrated, not general. Event verification, referring expressions, and temporal localization fall roughly 9 to 24 points at every model scale, while video question answering stays within 5.3 points of VideoMME and 2D spatial pointing shows no shortfall against BLINK. The temporal pillar is weakest in absolute terms: no system exceeds 55.7 mIoU on temporal localization or 37.3 SODA_c on dense video captioning. On tracking, frontier models come within roughly 5 points of specialist trackers over short horizons but separate as the horizon extends. Open-weight models lead 2D object localization outright, so neither scale nor proprietary access explains the pattern. Data, evaluation harness, and leaderboard: https://vantage-bench.org/

cs.CV cs.AI