cs.CV 2403.00476

TempCompass: Do Video LLMs Really Understand Videos?

TempCompass benchmark evaluates 8 SOTA Video LLMs across five temporal dimensions, revealing their poor temporal perception abilities, with average accuracy around 33.9%.

Yuanxin Liu, Shicheng Li, Yi Liu et al.

2024-03-01 367 citations 58
cs.LG 2402.19464

Curiosity-driven Red-teaming for Large Language Models

Curiosity-driven red teaming (CRT) leverages exploration rewards to enhance test coverage, successfully eliciting toxic responses from LLaMA2, with a 19.6% toxicity rate compared to 10.2% by baseline methods.

Zhang-Wei Hong, Idan Shenfeld, Tsun-Hsuan Wang et al.

2024-03-01 96 citations 34
cs.RO 2402.18319

A Multimodal Handover Failure Detection Dataset and Baselines

Proposed a multimodal handover failure detection dataset using video, force-torque, and gripper data; baseline methods include 3D CNN and action segmentation, achieving 67.9% accuracy.

Santosh Thoduka, Nico Hochgeschwender, Juergen Gall et al.

2024-02-28 33