cs.LG 2001.08361

Scaling Laws for Neural Language Models

This study formulates power-law scaling laws for neural language models, linking model size, data, and compute to performance, guiding efficient training strategies.

Jared Kaplan, Sam McCandlish, Tom Henighan et al.

2020-01-23 28
cs.LG 2001.07457

Learning to Control PDEs with Differentiable Physics

Hierarchical predictor-corrector framework with differentiable PDE solver enables long-term control of nonlinear PDE systems like Navier-Stokes.

Philipp Holl, Vladlen Koltun, Nils Thuerey

2020-01-21 34
cs.LG 2001.04451

Reformer: The Efficient Transformer

Reformer combines reversible residual layers and locality-sensitive hashing (LSH) attention, reducing complexity from O(L^2) to O(L log L) for long sequences.

Nikita Kitaev, Łukasz Kaiser, Anselm Levskaya

2020-01-14 49
cs.LG 2001.03040

Deep Network Approximation for Smooth Functions

This paper establishes near-optimal approximation bounds for deep ReLU networks approximating smooth functions, with errors of O(N^{-2s/d}L^{-2s/d}) when width and depth are optimized.

Jianfeng Lu, Zuowei Shen, Haizhao Yang et al.

2020-01-09 317 citations 29
cs.LG 1912.04977

Advances and Open Problems in Federated Learning

Proposes federated learning algorithms with privacy and efficiency improvements; experiments show 20% accuracy gain on CIFAR-10.

Peter Kairouz, H. Brendan McMahan, Brendan Avent et al.

2019-12-11 22
cs.LG 1912.02292

Deep Double Descent: Where Bigger Models and More Data Hurt

Introduces ‘Effective Model Complexity’ to explain double descent phenomena, revealing non-monotonic effects of model size and training epochs on generalization.

Preetum Nakkiran, Gal Kaplun, Yamini Bansal et al.

2019-12-05 1184 citations 38
cs.LG 1912.01603

Dream to Control: Learning Behaviors by Latent Imagination

Dreamer leverages latent space imagination with deep models and analytic gradients to achieve efficient long-horizon control on 20 visual tasks, surpassing prior methods.

Danijar Hafner, Timothy Lillicrap, Jimmy Ba et al.

2019-12-04 57
cs.LG 1912.00967

Continuous Graph Neural Networks

Proposes Continuous Graph Neural Networks (CGNN) using neural ODEs, addressing over-smoothing and capturing long-range dependencies, achieving state-of-the-art node classification accuracy.

Louis-Pascal A. C. Xhonneux, Meng Qu, Jian Tang

2019-12-03 210 citations 24