cs.LG 2402.10240

A Dynamical View of the Question of Why

Proposes a causal reasoning framework based on reinforcement learning, utilizing two key lemmas to quantify causality in multivariate stochastic processes.

Mehdi Fatemi, Sindhu Gowda

2024-02-15 45
cs.CL 2402.09353

DoRA: Weight-Decomposed Low-Rank Adaptation

DoRA enhances LoRA's learning capacity via weight decomposition, outperforming LoRA on multiple tasks.

Shih-Yang Liu, Chien-Yi Wang, Hongxu Yin et al.

2024-02-15 37
cs.LG 2402.08667

Target Score Matching

Introduces Target Score Matching (TSM), leveraging known target scores to improve low-noise score estimates, enhancing diffusion models’ accuracy.

Valentin De Bortoli, Michael Hutchinson, Peter Wirnsberger et al.

2024-02-14 40
cs.LG 2402.07871

Scaling Laws for Fine-Grained Mixture of Experts

Introduces granular hyperparameter G and scaling laws for MoE, optimizing training to outperform dense Transformers under various compute budgets.

Jakub Krajewski, Jan Ludziejewski, Kamil Adamczewski et al.

2024-02-13 32
cs.LG 2402.09470

Rolling Diffusion Models

Proposes Rolling Diffusion with sliding window scheduling, outperforming standard diffusion in complex temporal data, validated on Kinetics-600 and fluid simulations.

David Ruhe, Jonathan Heek, Tim Salimans et al.

2024-02-12 56