cs.LG 1703.00887

How to Escape Saddle Points Efficiently

Perturbed gradient descent efficiently escapes saddle points, converging to approximate second-order stationary points with nearly dimension-free complexity.

Chi Jin, Rong Ge, Praneeth Netrapalli et al.

2017-03-03 44
cs.LG 1702.08165

Reinforcement Learning with Deep Energy-Based Policies

Proposes deep energy-based soft Q-learning for continuous spaces, improving exploration and skill transfer with a Boltzmann policy approximation.

Tuomas Haarnoja, Haoran Tang, Pieter Abbeel et al.

2017-02-27 1640 citations 33
cs.LG 1702.07539

Tight Bounds for Bandit Combinatorial Optimization

This paper proves the optimal regret in combinatorial bandits grows as \(\widetilde{\Theta}(k^{3/2}\sqrt{dT})\), refuting prior conjectures.

Alon Cohen, Tamir Hazan, Tomer Koren

2017-02-24 38
cs.LG 1610.02995

Extrapolation and learning equations

Proposes Equation Learner (EQL), a neural network that learns analytical expressions and excels in extrapolation beyond training data.

Georg Martius, Christoph H. Lampert

2016-10-11 43