cs.CL 2005.14165

Language Models are Few-Shot Learners

Scaling up to 175B parameters, GPT-3 achieves remarkable few-shot and zero-shot NLP performance, surpassing many fine-tuned models.

Tom B. Brown, Benjamin Mann, Nick Ryder et al.

2020-05-29 37
cs.LG 2005.13590

Demystifying Orthogonal Monte Carlo and Beyond

This paper advances Orthogonal Monte Carlo (OMC) theory by applying negative dependence, deriving exponential error bounds, and introduces Near-Orthogonal Monte Carlo (NOMC) for superior high-dimensional sampling.

Han Lin, Haoxian Chen, Tianyi Zhang et al.

2020-05-28 11 citations 36
cs.RO 2005.12813

AlphaPilot: Autonomous Drone Racing

Combining learned data abstraction, nonlinear filtering, and time-optimal trajectory planning, the autonomous drone racing system achieved second place at the 2019 AlphaPilot Challenge, reaching speeds up to 8 m/s.

Philipp Foehn, Dario Brescianini, Elia Kaufmann et al.

2020-05-26 32
cs.LG 2005.09841

Best Arm Identification in Spectral Bandits

Proposes a gradient ascent-based optimal sampling strategy for spectral bandit best-arm identification under graph smoothness constraints, achieving asymptotic optimality.

Tomáš Kocák, Aurélien Garivier

2020-05-20 46
cs.IR 2005.09683

Neural Collaborative Filtering vs. Matrix Factorization Revisited

This study compares neural collaborative filtering (MLP) with matrix factorization (dot product), showing that after hyperparameter tuning, dot product outperforms MLP in recommendation tasks with lower complexity.

Steffen Rendle, Walid Krichene, Li Zhang et al.

2020-05-20 26
cs.CL 2005.07683

Movement Pruning: Adaptive Sparsity by Fine-Tuning

Movement pruning leverages first-order importance scores during fine-tuning, achieving high sparsity with minimal accuracy loss, outperforming magnitude pruning.

Victor Sanh, Thomas Wolf, Alexander M. Rush

2020-05-16 48
cs.LG 2005.05951

MOReL : Model-Based Offline Reinforcement Learning

MOReL is a model-based offline reinforcement learning algorithm that achieves state-of-the-art performance on the D4RL benchmark.

Rahul Kidambi, Aravind Rajeswaran, Praneeth Netrapalli et al.

2020-05-13 3