cs.LG 2006.10598

Neural Parameter Allocation Search

Proposes NPAS and SSNs for automatic parameter-efficient neural network training under arbitrary budgets.

Bryan A. Plummer, Nikoli Dryden, Julius Frost et al.

2020-06-18 38
cs.LG 2006.10029

Big Self-Supervised Models are Strong Semi-Supervised Learners

This paper introduces semi-supervised learning with Big Self-Supervised Models (SimCLRv2), achieving 73.9% top-1 accuracy on ImageNet with only 1% labels, outperforming previous methods by 10x.

Ting Chen, Simon Kornblith, Kevin Swersky et al.

2020-06-18 2601 citations 40
cs.LG 2006.09503

Memory-Efficient Pipeline-Parallel DNN Training

Proposes PipeDream-2BW, a memory-efficient pipeline training system for large-scale DNNs, achieving up to 20× speedup via double-buffered weight updates and automated model partitioning.

Deepak Narayanan, Amar Phanishayee, Kaiyu Shi et al.

2020-06-17 307 citations 42
cs.LG 2006.04768

Linformer: Self-Attention with Linear Complexity

Linformer reduces self-attention complexity from O(n²) to O(n) using low-rank approximation, enabling efficient long-sequence modeling.

Sinong Wang, Belinda Z. Li, Madian Khabsa et al.

2020-06-09 46
cs.LG 2005.13590

Demystifying Orthogonal Monte Carlo and Beyond

This paper advances Orthogonal Monte Carlo (OMC) theory by applying negative dependence, deriving exponential error bounds, and introduces Near-Orthogonal Monte Carlo (NOMC) for superior high-dimensional sampling.

Han Lin, Haoxian Chen, Tianyi Zhang et al.

2020-05-28 11 citations 35