LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models
LLaMA-VID introduces a dual-token strategy enabling efficient long video understanding, surpassing previous methods.
Yanwei Li, Chengyao Wang, Jiaya Jia
LLaMA-VID introduces a dual-token strategy enabling efficient long video understanding, surpassing previous methods.
Yanwei Li, Chengyao Wang, Jiaya Jia
Proposes MVBench, a comprehensive benchmark for evaluating temporal understanding in multi-modal video models, with a novel static-to-dynamic task transformation and automated QA generation; VideoChat2 surpasses SOTA by over 15%.
Kunchang Li, Yali Wang, Yinan He et al.
Proposes a cross-view dense video captioning framework using adversarial learning to transfer knowledge from web instructional videos to egocentric videos.
Takehiko Ohkawa, Takuma Yagi, Taichi Nishimura et al.
Proposes CCoT, a zero-shot chain-of-thought method using scene graphs to improve large multimodal models' compositional reasoning, boosting benchmark scores.
Chancharik Mitra, Brandon Huang, Trevor Darrell et al.
MagicAnimate uses diffusion models with temporal attention and appearance encoders to produce long, coherent human animations, outperforming baselines by over 38%.
Zhongcong Xu, Jianfeng Zhang, Jun Hao Liew et al.
SeeSR leverages degradation-aware semantic prompts to guide diffusion models, achieving realistic super-resolution with strong semantic fidelity.
Rongyuan Wu, Tao Yang, Lingchen Sun et al.
Mip-Splatting introduces 3D frequency-constrained smoothing and Mip filtering to eliminate aliasing artifacts in Gaussian Splatting, improving multi-scale view synthesis.
Zehao Yu, Anpei Chen, Binbin Huang et al.
MeshGPT uses decoder-only Transformer with quantized geometric vocabulary to generate compact, sharp 3D meshes, improving shape coverage by 9%.
Yawar Siddiqui, Antonio Alliegro, Alexey Artemov et al.
Proposes Stable Video Diffusion with a three-stage training strategy, leveraging large-scale curated datasets to significantly improve high-resolution video generation.
Andreas Blattmann, Tim Dockhorn, Sumith Kulal et al.
AutoEval-Video is a benchmark for evaluating large vision-language models in video QA; GPT-4V excels but lags behind human accuracy of 72.8%.
Xiuyuan Chen, Yuan Lin, Yuchen Zhang et al.
This study systematically evaluates GPT-4V's capabilities in geographic and geospatial tasks via a custom benchmark, revealing strengths in fine-grained recognition and spatial reasoning.
Jonathan Roberts, Timo Lüddecke, Rehan Sheikh et al.
ZeroPS transfers 2D pretrained models SAM and GLIP to zero-shot 3D part segmentation without fine-tuning, based on multi-view relations.
Yuheng Xue, Nenglun Chen, Jun Liu et al.
Introduces ShareGPT4V dataset with 1.2M detailed image captions, significantly boosting large multimodal model performance, achieving multiple SOTA results.
Lin Chen, Jinsong Li, Xiaoyi Dong et al.
Proposes a dual-scale ordinal shading framework for high-res intrinsic image decomposition, outperforming SOTA with detailed, globally consistent results.
Chris Careaga, Yağız Aksoy
GLAD uses multi-scale view alignment and background debiasing to improve unsupervised video domain adaptation with large domain gaps.
Hyogun Lee, Kyungho Bae, Seong Jong Ha et al.
Proposes synchronized multi-view diffusion for consistent text-guided 3D texturing, improving coherence by sharing latent content during denoising.
Yuxin Liu, Minshan Xie, Hanyuan Liu et al.
Introduces Hierarchical Relation Head and Commonsense Validation, boosting scene graph accuracy by over 10% R@50.
Bowen Jiang, Zhijun Zhuang, Shreyas S. Shivakumar et al.
PF-LRM is a pose-free 3D reconstruction model achieving joint shape and pose prediction in 1.3 seconds using transformer-based token interaction and point cloud supervision.
Peng Wang, Hao Tan, Sai Bi et al.
AutoStory combines LLM-based layout planning with diffusion models to generate diverse, high-quality storytelling images from minimal input, reducing manual effort.
Wen Wang, Canyu Zhao, Hao Chen et al.
LEO, a multi-modal embodied agent trained via 3D VL alignment and instruction tuning, significantly advances 3D scene understanding and interaction.
Jiangyong Huang, Silong Yong, Xiaojian Ma et al.