cs.LG 1502.05477

Trust Region Policy Optimization

TRPO (Trust Region Policy Optimization) guarantees monotonic policy improvement using KL constraints, excelling in large neural network policy training for robotics and Atari games.

John Schulman, Sergey Levine, Philipp Moritz et al.

2015-02-19 8275 citations 49
cs.LG 1502.04681

Unsupervised Learning of Video Representations using LSTMs

Proposes a multi-layer LSTM encoder-decoder framework for unsupervised video representation learning, improving action recognition especially with limited labeled data.

Nitish Srivastava, Elman Mansimov, Ruslan Salakhutdinov

2015-02-17 58
cs.SI 1501.06247

Reciprocal Recommendation System for Online Dating

Proposed reciprocal scoring-based online dating recommendation leveraging multi-dimensional similarity measures, improving matching accuracy and user satisfaction.

Peng Xia, Benyuan Liu, Yizhou Sun et al.

2015-01-26 58
cs.LG 1412.7149

Deep Fried Convnets

Deep Fried ConvNets replace fully connected layers with adaptive Fastfood transforms, reducing parameters by over 90% without accuracy loss.

Zichao Yang, Marcin Moczulski, Misha Denil et al.

2014-12-23 56
cs.LG 1412.6980

Adam: A Method for Stochastic Optimization

Adam combines first and second moment estimates for adaptive learning rates, accelerating large-scale stochastic optimization.

Diederik P. Kingma, Jimmy Ba

2014-12-22 37
cs.LG 1412.6651

Deep learning with Elastic Averaging SGD

Elastic Averaging SGD (EASGD) leverages elastic force to enable efficient distributed deep learning, reducing communication overhead while enhancing exploration, leading to faster convergence and better accuracy.

Sixin Zhang, Anna Choromanska, Yann LeCun

2014-12-20 646 citations 38
cs.LG 1412.0233

The Loss Surfaces of Multilayer Networks

A spherical spin-glass mapping explains why SGD in large networks reaches low-energy, high-quality local minima.

Anna Choromanska, Mikael Henaff, Michael Mathieu et al.

2014-11-30 17