cs.LG 1412.0233

The Loss Surfaces of Multilayer Networks

A spherical spin-glass mapping explains why SGD in large networks reaches low-energy, high-quality local minima.

Anna Choromanska, Mikael Henaff, Michael Mathieu et al.

2014-11-30 17
cs.LG 1410.1141

On the Computational Efficiency of Training Neural Networks

This paper analyzes the computational complexity of training neural networks, showing over-parameterized networks are easier to optimize and proposing polynomial activation-based algorithms for depth-2 and depth-3 networks.

Roi Livni, Shai Shalev-Shwartz, Ohad Shamir

2014-10-05 51
cs.LG 1306.0686

Online Learning under Delayed Feedback

Proposes BOLD and QPM-D algorithms for online learning with delayed feedback; theoretical bounds show multiplicative regret in adversarial and additive in stochastic settings.

Pooria Joulani, András György, Csaba Szepesvári

2013-06-04 20
cs.LG 1206.6471

On Causal and Anticausal Learning

Explores causal learning's impact on semi-supervised learning with hypotheses and validation.

Bernhard Schoelkopf, Dominik Janzing, Jonas Peters et al.

2012-06-28 2
cs.LG 1206.4655

Modelling transition dynamics in MDPs with RKHS embeddings

Proposes a nonparametric RKHS embedding approach for modeling MDP transition dynamics, outperforming Gaussian processes and NPDP in efficiency and accuracy.

Steffen Grunewalder, Guy Lever, Luca Baldassarre et al.

2012-06-18 71