Generative Multimodal Models are In-Context Learners
Emu2, with 3.7 billion parameters, significantly enhances few-shot and zero-shot multimodal understanding and generation capabilities.
Quan Sun, Yufeng Cui, Xiaosong Zhang et al.
Emu2, with 3.7 billion parameters, significantly enhances few-shot and zero-shot multimodal understanding and generation capabilities.
Quan Sun, Yufeng Cui, Xiaosong Zhang et al.
pixelSplat employs probabilistic 3D Gaussian primitives with reparameterization for scalable, real-time scene reconstruction, outperforming SOTA with 2.5× faster rendering and explicit 3D representations.
David Charatan, Sizhe Li, Andrea Tagliasacchi et al.
Pose2Gaze combines CNN and ST-GCN to predict eye gaze from full-body poses, reducing mean angular error by 24% across datasets.
Zhiming Hu, Jiahui Xu, Syn Schmitt et al.
DriveMLM aligns multimodal LLMs with behavioral states, enabling closed-loop autonomous driving with 3.2-4.7 point improvements on CARLA.
Erfei Cui, Wenhai Wang, Zhiqi Li et al.
Proposes 3D Gaussian Splatting for animatable avatars, training in 30 min, real-time at 50+ FPS, outperforming SOTA in speed and quality.
Zhiyin Qian, Shaofei Wang, Marko Mihajlovic et al.
Holodeck combines GPT-4 and Objaverse assets to generate diverse, semantically accurate 3D environments from text prompts, advancing scene diversity and layout realism.
Yue Yang, Fan-Yun Sun, Luca Weihs et al.
CogAgent is an 18-billion-parameter visual language model that excels in GUI understanding and navigation, achieving state-of-the-art results on multiple VQA benchmarks.
Wenyi Hong, Weihan Wang, Qingsong Lv et al.
Proposes KMCGs-enhanced diffusion-based motion style transfer, supporting ten dance styles with improved content preservation and scalability.
Wenjie Yin, Yi Yu, Hang Yin et al.
Proposes CoPoNeRF, a unified framework for pose-free novel view synthesis with significant improvements over prior methods, achieving high-quality rendering without explicit camera poses.
Sunghwan Hong, Jaewoo Jung, Heeseong Shin et al.
W.A.L.T employs causal encoding and window attention to achieve state-of-the-art photorealistic video generation at 512×896 resolution, 8 fps.
Agrim Gupta, Lijun Yu, Kihyuk Sohn et al.
Upscale-A-Video leverages latent diffusion with local-global strategies for temporally consistent video super-resolution, surpassing existing methods with detailed results.
Shangchen Zhou, Peiqing Yang, Jianyi Wang et al.
ASH uses 2D texture space parameterization of Gaussian splats for real-time high-fidelity animated human rendering.
Haokai Pang, Heming Zhu, Adam Kortylewski et al.
DiffMatte employs diffusion models for multi-step iterative natural image matting, reducing SAD by 8%.
Yihan Hu, Yiheng Lin, Wei Wang et al.
HaMeR employs a Transformer-based architecture with large-scale data, achieving state-of-the-art 3D hand mesh reconstruction from monocular images, outperforming previous methods.
Georgios Pavlakos, Dandan Shan, Ilija Radosavovic et al.
RAVE employs noise shuffling in diffusion sampling for fast, zero-shot, temporally consistent video editing.
Ozgur Kara, Bariscan Kurtkaya, Hidir Yesiltepe et al.
DreamVideo employs textual inversion and lightweight adapters to enable customizable video synthesis, modeling subject appearance and motion separately with high efficiency.
Yujie Wei, Shiwei Zhang, Zhiwu Qing et al.
Proposes Representation-Conditioned Generation (RCG) framework using self-supervised encoder-generated semantic representations to greatly improve unconditional image generation, achieving a new SOTA FID of 2.15.
Tianhong Li, Dina Katabi, Kaiming He
OneLLM aligns eight modalities with language using a unified framework, enhancing multimodal understanding.
Jiaming Han, Kaixiong Gong, Yiyuan Zhang et al.
MotionCtrl introduces a dual-module framework for independent control of camera and object motion in video generation, leveraging diffusion models and trajectory-based conditioning.
Zhouxia Wang, Ziyang Yuan, Xintao Wang et al.
VPD enhances vision-language models by distilling programmatic reasoning, surpassing existing models.
Yushi Hu, Otilia Stretcu, Chun-Ta Lu et al.