cs.LG 1502.04681

Unsupervised Learning of Video Representations using LSTMs

Proposes a multi-layer LSTM encoder-decoder framework for unsupervised video representation learning, improving action recognition especially with limited labeled data.

Nitish Srivastava, Elman Mansimov, Ruslan Salakhutdinov

2015-02-17 56
cs.LG 1412.7149

Deep Fried Convnets

Deep Fried ConvNets replace fully connected layers with adaptive Fastfood transforms, reducing parameters by over 90% without accuracy loss.

Zichao Yang, Marcin Moczulski, Misha Denil et al.

2014-12-23 54
cs.LG 1412.6980

Adam: A Method for Stochastic Optimization

Adam combines first and second moment estimates for adaptive learning rates, accelerating large-scale stochastic optimization.

Diederik P. Kingma, Jimmy Ba

2014-12-22 36
cs.LG 1412.6651

Deep learning with Elastic Averaging SGD

Elastic Averaging SGD (EASGD) leverages elastic force to enable efficient distributed deep learning, reducing communication overhead while enhancing exploration, leading to faster convergence and better accuracy.

Sixin Zhang, Anna Choromanska, Yann LeCun

2014-12-20 646 citations 38
cs.LG 1412.0233

The Loss Surfaces of Multilayer Networks

A spherical spin-glass mapping explains why SGD in large networks reaches low-energy, high-quality local minima.

Anna Choromanska, Mikael Henaff, Michael Mathieu et al.

2014-11-30 15
cs.LG 1410.1141

On the Computational Efficiency of Training Neural Networks

This paper analyzes the computational complexity of training neural networks, showing over-parameterized networks are easier to optimize and proposing polynomial activation-based algorithms for depth-2 and depth-3 networks.

Roi Livni, Shai Shalev-Shwartz, Ohad Shamir

2014-10-05 49
cs.LG 1306.0686

Online Learning under Delayed Feedback

Proposes BOLD and QPM-D algorithms for online learning with delayed feedback; theoretical bounds show multiplicative regret in adversarial and additive in stochastic settings.

Pooria Joulani, András György, Csaba Szepesvári

2013-06-04 20