cs.CV 2403.00476

TempCompass: Do Video LLMs Really Understand Videos?

TempCompass benchmark evaluates 8 SOTA Video LLMs across five temporal dimensions, revealing their poor temporal perception abilities, with average accuracy around 33.9%.

Yuanxin Liu, Shicheng Li, Yi Liu et al.

2024-03-01 367 citations 54
cs.CV 2402.10099

Any-Shift Prompting for Generalization over Distributions

Proposes Any-Shift Prompting, a hierarchical Bayesian framework, achieving significant improvements over 23 datasets in cross-distribution generalization.

Zehao Xiao, Jiayi Shen, Mohammad Mahdi Derakhshani et al.

2024-02-16 27 citations 31