MoVA: Adapting Mixture of Vision Experts to Multimodal Context
MoVA adapts a mixture of vision experts to multimodal contexts, enhancing performance through a coarse-to-fine mechanism.
Zhuofan Zong, Bingqi Ma, Dazhong Shen et al.
MoVA adapts a mixture of vision experts to multimodal contexts, enhancing performance through a coarse-to-fine mechanism.
Zhuofan Zong, Bingqi Ma, Dazhong Shen et al.
PhysDreamer leverages video generation priors to estimate material fields, enabling realistic interactive 3D object dynamics without real-world material data.
Tianyuan Zhang, Hong-Xing Yu, Rundi Wu et al.
This paper analyzes the content bias in Fréchet Video Distance (FVD), revealing its overemphasis on frame quality over motion continuity, rooted in feature extractor biases.
Songwei Ge, Aniruddha Mahapatra, Gaurav Parmar et al.
Blink benchmark evaluates 14 core visual perception tasks using multimodal models like GPT-4V and Gemini, achieving only 51.26% accuracy, far below human performance of 95.70%.
Xingyu Fu, Yushi Hu, Bangzheng Li et al.
The paper introduces explicit 3D-feature viewpoint control for customized diffusion generation, but the supplied text reports no numerical metrics.
Nupur Kumari, Grace Su, Richard Zhang et al.
Proposes LAPTOP-Diff, combining layer pruning and normalized distillation to compress diffusion models with only 4% performance loss at 50% pruning.
Dingkun Zhang, Sijia Li, Chen Chen et al.
Semantics-Aware Attention Guidance (SAG) improves cancer diagnosis accuracy on whole slide images by integrating tissue and cellular cues, boosting performance by over 6%.
Kechun Liu, Wenjun Wu, Joann G. Elmore et al.
MK-SGN combines multimodal fusion and knowledge distillation for skeleton-based action recognition, achieving 98% energy reduction.
Naichuan Zheng, Hailun Xia, Zeyu Liang et al.
Introduces MMInA benchmark for evaluating multihop multimodal web agents, using 1050 real evolving websites to assess long-range reasoning.
Shulin Tian, Ziniu Zhang, Liangyu Chen et al.
Ctrl-Adapter efficiently adapts ControlNet, achieving superior performance on COCO and DAVIS 2017.
Han Lin, Jaemin Cho, Abhay Zala et al.
SparseOcc employs sparse latent representations for semantic occupancy prediction, reducing FLOPs by 74.9% and increasing mIoU from 12.8% to 14.1%.
Pin Tang, Zhongdao Wang, Guoqing Wang et al.
This study probes large-scale visual models' 3D awareness via task-specific probes, revealing significant limitations in depth, normals, and multiview consistency.
Mohamed El Banani, Amit Raj, Kevis-Kokitsi Maninis et al.
RealmDreamer leverages inpainting and depth diffusion for text-driven 3D scene synthesis, achieving 95.5% user preference in evaluations.
Jaidev Shriram, Alex Trevithick, Lingjie Liu et al.
InternLM-XComposer2-4KHD supports dynamic resolutions from 336 pixels to 4K, significantly enhancing high-resolution vision-language understanding.
Xiaoyi Dong, Pan Zhang, Yuhang Zang et al.
Ferret-UI enhances mobile UI understanding using multimodal LLMs, surpassing GPT-4V.
Keen You, Haotian Zhang, Eldon Schoop et al.
VAR employs multi-scale autoregressive prediction, reducing FID to 1.73 and increasing speed 20×, surpassing diffusion models.
Keyu Tian, Yi Jiang, Zehuan Yuan et al.
Proposes neural implicit-based method for reconstructing digital twins of unknown multi-part articulated objects from two RGB-D scans, outperforming prior approaches.
Yijia Weng, Bowen Wen, Jonathan Tremblay et al.
Proposes a multi-label contrastive learning framework for style descriptors, achieving state-of-the-art style retrieval accuracy.
Gowthami Somepalli, Anubhav Gupta, Kamal Gupta et al.
VQAScore leverages VQA models to evaluate text-image alignment, surpassing CLIPScore with state-of-the-art results on 8 benchmarks.
Zhiqiu Lin, Deepak Pathak, Baiqi Li et al.
EvLight leverages multi-scale fusion and SNR-guided feature selection, achieving +1.14dB PSNR over frame-based methods on large real-world event-image datasets.
Guoqiang Liang, Kanghao Chen, Hangyu Li et al.