cs.CV 2404.12391

On the Content Bias in Fréchet Video Distance

This paper analyzes the content bias in Fréchet Video Distance (FVD), revealing its overemphasis on frame quality over motion continuity, rooted in feature extractor biases.

Songwei Ge, Aniruddha Mahapatra, Gaurav Parmar et al.

2024-04-19 62 citations 42
cs.CV 2404.08636

Probing the 3D Awareness of Visual Foundation Models

This study probes large-scale visual models' 3D awareness via task-specific probes, revealing significant limitations in depth, normals, and multiview consistency.

Mohamed El Banani, Amit Raj, Kevis-Kokitsi Maninis et al.

2024-04-13 44
cs.CV 2404.01292

Measuring Style Similarity in Diffusion Models

Proposes a multi-label contrastive learning framework for style descriptors, achieving state-of-the-art style retrieval accuracy.

Gowthami Somepalli, Anubhav Gupta, Kamal Gupta et al.

2024-04-02 44