cs.CV 2606.17027

MeshLoom: Feed-Forward Non-Rigid Registration of Mesh Sequences

MeshLoom is a feed-forward non-rigid mesh registration network that reconstructs vertex deformations across sequences within seconds, outperforming state-of-the-art methods.

Jianqi Chen, Jiraphon Yenphraphai, Xiangjun Tang et al.

2026-06-16 356
cs.CV 2606.15389

Timestep Rescheduling in Diffusion Inversion

Proposes a non-uniform timestep rescheduling method for diffusion inversion, reducing errors and improving image reconstruction accuracy by leveraging error analysis and dynamic programming.

Shangquan Sun, Ting Gong, Zhirui Liu et al.

2026-06-14 45
cs.CV 2606.15341

CausalDrive: Real-time Causal World Models for Autonomous Driving

CausalDrive employs a real-time causal autoregressive world model with flow-matching and self-distillation, achieving 12 FPS interactive autonomous driving simulation without future layout conditioning.

Tianyi Yan, Huan Zheng, Dubing Chen et al.

2026-06-13 50
cs.CV 2606.14958

MVEB: Massive Video Embedding Benchmark

MVEB benchmarks 23 tasks across 33 models, revealing diverse strengths and limitations in multi-modal video embeddings.

Adnan El Assadi, Roman Solomatin, Isaac Chung et al.

2026-06-13 74
cs.CV 2606.14703

Gaze Heads: How VLMs Look at What They Describe

This study identifies a small set of attention heads—gaze heads—in VLMs that causally track the current description region, enabling effective inference-time control via attention masks.

Rohit Gandikota, David Bau

2026-06-13 215