cs.LG 2305.13301

Training Diffusion Models with Reinforcement Learning

Introduces DDPO, a reinforcement learning framework that directly optimizes diffusion models for target-specific objectives, improving image quality and alignment.

Kevin Black, Michael Janner, Yilun Du et al.

2023-05-23 38
cs.CV 2305.13077

ControlVideo: Training-free Controllable Text-to-Video Generation

ControlVideo employs full cross-frame attention and hierarchical sampling to enable training-free, high-quality controllable text-to-video generation, outperforming state-of-the-art methods.

Yabo Zhang, Yuxiang Wei, Dongsheng Jiang et al.

2023-05-22 379 citations 44
cs.CL 2305.12535

Explaining How Transformers Use Context to Build Predictions

Proposes ALTI-Logit, a residual and attention-based explanation method outperforming gradients and perturbations, with strong results on linguistic and translation tasks.

Javier Ferrando, Gerard I. Gállego, Ioannis Tsiamas et al.

2023-05-22 52
eess.AS 2305.11834

Pengi: An Audio Language Model for Audio Tasks

Pengi unifies audio tasks as text generation, reaching 0.4667 SPIDEr on AudioCaps and 0.9195 accuracy on ESC50.

Soham Deshmukh, Benjamin Elizalde, Rita Singh et al.

2023-05-20 28