cs.CV 2406.08035

LVBench: An Extreme Long Video Understanding Benchmark

LVBench evaluates long video understanding, focusing on long-term memory and reasoning, with a dataset averaging 4101 seconds, covering six core tasks.

Weihan Wang, Zehai He, Wenyi Hong et al.

2024-06-12 43
cs.CV 2406.04264

MLVU: Benchmarking Multi-task Long Video Understanding

MLVU benchmark evaluates 23 multimodal models on long videos (up to 2 hours) across 9 tasks, revealing significant room for improvement in long video understanding.

Junjie Zhou, Yan Shu, Bo Zhao et al.

2024-06-07 262 citations 35
cs.CV 2406.02507

Guiding a Diffusion Model with a Bad Version of Itself

Proposes self-guidance using weaker models to enhance diffusion model image quality, achieving a record FID of 1.25 on ImageNet-512.

Tero Karras, Miika Aittala, Tuomas Kynkäänniemi et al.

2024-06-05 319 citations 38
cs.CV 2406.00885

Visual place recognition for aerial imagery: A survey

This paper introduces a novel evaluation framework for aerial Visual Place Recognition (VPR), emphasizing multi-scale map construction and overlap optimization.

Ivan Moskalenko, Anastasiia Kornilova, Gonzalo Ferrer

2024-06-03 36