cs.CV 2308.13712

Residual Denoising Diffusion Models

Residual Denoising Diffusion Model (RDDM) introduces dual diffusion processes for unified image generation and restoration, leveraging residuals and noise with independent scheduling.

Jiawei Liu, Qiang Wang, Huijie Fan et al.

2023-08-26 43
cs.CV 2308.07498

DREAMWALKER: Mental Planning for Continuous Vision-Language Navigation

DREAMWALKER employs explicit world models with environment graphs and scene synthesizers, combined with Monte Carlo Tree Search, to enable strategic mental planning in continuous vision-language navigation, achieving over 78% success rate.

Hanqing Wang, Wei Liang, Luc Van Gool et al.

2023-08-15 121 citations 32
cs.CV 2308.06571

ModelScope Text-to-Video Technical Report

ModelScopeT2V uses diffusion with spatio-temporal blocks, 1.7B parameters, achieving superior text-to-video synthesis with high temporal coherence.

Jiuniu Wang, Hangjie Yuan, Dayou Chen et al.

2023-08-12 34