cs.CV 2606.14958

MVEB: Massive Video Embedding Benchmark

MVEB benchmarks 23 tasks across 33 models, revealing diverse strengths and limitations in multi-modal video embeddings.

Adnan El Assadi, Roman Solomatin, Isaac Chung et al.

2026-06-13 75
cs.CV 2606.14703

Gaze Heads: How VLMs Look at What They Describe

This study identifies a small set of attention heads—gaze heads—in VLMs that causally track the current description region, enabling effective inference-time control via attention masks.

Rohit Gandikota, David Bau

2026-06-13 216
cs.CL 2606.14626

Characterizing Cultural Localization in AI-Generated Stories

Proposes a method combining lexical token analysis and multi-word similarity to quantify cultural localization in AI-generated stories, revealing only 9-17% of vocabulary accounts for cultural differences.

Shaily Bhatt, Supriti Vijay, Jeremiah Milbauer et al.

2026-06-13 237