cs.CV 2312.02696

Analyzing and Improving the Training Dynamics of Diffusion Models

This paper introduces layer magnitude preservation and post-hoc EMA tuning, reducing ImageNet-512 FID from 2.41 to 1.81, significantly improving diffusion model training stability and quality.

Tero Karras, Miika Aittala, Jaakko Lehtinen et al.

2023-12-05 32
cs.CV 2312.00858

DeepCache: Accelerating Diffusion Models for Free

DeepCache accelerates diffusion models by caching features, achieving 2.3× speedup on Stable Diffusion v1.5 with minimal quality loss without retraining.

Xinyin Ma, Gongfan Fang, Xinchao Wang

2023-12-02 35
cs.CV 2311.18836

ChatPose: Chatting about 3D Human Pose

ChatPose employs multimodal LLMs with embedded SMPL parameters for 3D human pose understanding and generation.

Yao Feng, Jing Lin, Sai Kumar Dwivedi et al.

2023-12-01 42