2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining
Introduces a multimodal textbook dataset to enhance VLMs' performance in knowledge reasoning tasks.
Wenqi Zhang, Hang Zhang, Xin Li et al.
Introduces a multimodal textbook dataset to enhance VLMs' performance in knowledge reasoning tasks.
Wenqi Zhang, Hang Zhang, Xin Li et al.
This paper introduces LS-GAN, a latent-space GAN for human motion synthesis, achieving FID 0.482 with 91% FLOPs reduction compared to diffusion models.
Avinash Amballa, Gayathri Akkinapalli, Vinitra Muralikrishnan
Survey on visual grounding, analyzing new concepts like grounded pre-training and multimodal LLMs, providing comprehensive research directions.
Linhui Xiao, Xiaoshan Yang, Xiangyuan Lan et al.
DrivingGPT employs a multimodal autoregressive Transformer to unify world modeling and path planning, outperforming baselines with superior video quality and planning scores.
Yuntao Chen, Yuqi Wang, Zhaoxiang Zhang
HSfM combines deep learning and SfM to jointly reconstruct multiple human meshes, scene point clouds, and camera parameters with high accuracy.
Lea Müller, Hongsuk Choi, Anthony Zhang et al.
Introduces Causally Regularized Tokenization (CRT) to optimize visual token compression, enabling parameter and token reduction by half while maintaining state-of-the-art generation quality.
Vivek Ramanujan, Kushal Tirumala, Armen Aghajanyan et al.
Combining DreamBooth and contrastive learning, this method uses only 3 real images to generate synthetic data, significantly outperforming pretrained models across tasks.
Shobhita Sundaram, Julia Chae, Yonglong Tian et al.
LeviTor integrates depth estimation with K-means clustered control points to enable precise 3D trajectory control in image-to-video synthesis, outperforming prior 2D methods.
Hanlin Wang, Hao Ouyang, Qiuyu Wang et al.
Generative Multiview Relighting combines diffusion harmonization with NeRF-Casting, reaching 31.34 PSNR on Objaverse under extreme lighting variation.
Hadi Alzayer, Philipp Henzler, Jonathan T. Barron et al.
Prompt-A-Video leverages preference-aligned LLMs with reward-guided evolution, boosting video quality metrics by over 0.2 on average across models.
Yatai Ji, Jiacheng Zhang, Jie Wu et al.
Introduces ScaMo framework, validating that motion generation performance follows a logarithmic scaling law with compute, with model size and vocabulary size following power laws.
Shunlin Lu, Jingbo Wang, Zeyu Lu et al.
Introduced VSI-Bench, a video-based benchmark, revealing models achieve 45% accuracy in spatial reasoning, subhuman but improving with cognitive map generation.
Jihan Yang, Shusheng Yang, Anjali W. Gupta et al.
VideoDPO employs OmniScore for multi-dimensional preference alignment in video diffusion, significantly improving visual fidelity and semantic accuracy.
Runtao Liu, Haoyu Wu, Zheng Ziqiang et al.
MetaMorph employs VPiT to enable a pretrained LLM for joint visual understanding and generation, achieving high performance with minimal data.
Shengbang Tong, David Fan, Jiachen Zhu et al.
InstructSeg unifies image and video instructed segmentation using multi-modal large language models, achieving state-of-the-art results with a single end-to-end framework.
Cong Wei, Yujie Zhong, Haoxian Tan et al.
CG-Bench is a clue-grounded QA benchmark for long video understanding with 12,129 QA pairs.
Guo Chen, Yicheng Liu, Yifei Huang et al.
DEFAME employs a six-stage dynamic multimodal evidence retrieval framework, surpassing traditional text-only methods with significant accuracy gains.
Tobias Braun, Mark Rothermel, Marcus Rohrbach et al.
Diffusion-based 3D scene reconstruction from a single RGB image, achieving 12.04% AP3D and 13.43% F-Score improvements.
Manuel Dahnert, Angela Dai, Norman Müller et al.
FlowEdit edits SD3/FLUX without inversion and reaches SOTA.
Vladimir Kulikov, Matan Kleiner, Inbar Huberman-Spiegelglas et al.
CogNav employs LLM-based cognitive modeling with a heterogeneous map, boosting ObjectNav success by at least 14%.
Yihan Cao, Jiazhao Zhang, Zhinan Yu et al.