cs.CV 2501.01449

LS-GAN: Human Motion Synthesis with Latent-space GANs

This paper introduces LS-GAN, a latent-space GAN for human motion synthesis, achieving FID 0.482 with 91% FLOPs reduction compared to diffusion models.

Avinash Amballa, Gayathri Akkinapalli, Vinitra Muralikrishnan

2024-12-30 11 citations 40
cs.CV 2412.20206

Toward Visual Grounding: A Survey

Survey on visual grounding, analyzing new concepts like grounded pre-training and multimodal LLMs, providing comprehensive research directions.

Linhui Xiao, Xiaoshan Yang, Xiangyuan Lan et al.

2024-12-29 31
cs.CV 2412.17806

Reconstructing People, Places, and Cameras

HSfM combines deep learning and SfM to jointly reconstruct multiple human meshes, scene point clouds, and camera parameters with high accuracy.

Lea Müller, Hongsuk Choi, Anthony Zhang et al.

2024-12-24 25
cs.CV 2412.16156

Personalized Representation from Personalized Generation

Combining DreamBooth and contrastive learning, this method uses only 3 real images to generate synthetic data, significantly outperforming pretrained models across tasks.

Shobhita Sundaram, Julia Chae, Yonglong Tian et al.

2024-12-21 34
cs.CV 2412.15214

LeviTor: 3D Trajectory Oriented Image-to-Video Synthesis

LeviTor integrates depth estimation with K-means clustered control points to enable precise 3D trajectory control in image-to-video synthesis, outperforming prior 2D methods.

Hanlin Wang, Hao Ouyang, Qiuyu Wang et al.

2024-12-20 33