cs.LG 2009.04416

Phasic Policy Gradient

PPG introduces phased training with separate policy and value updates, boosting sample efficiency by ~30% on Procgen benchmarks.

Karl Cobbe, Jacob Hilton, Oleg Klimov et al.

2020-09-10 30
cs.LG 2009.00236

A Survey of Deep Active Learning

Proposes DeepAL framework combining Bayesian sampling and batch strategies, reducing labeling costs by 30% while maintaining high accuracy.

Pengzhen Ren, Yun Xiao, Xiaojun Chang et al.

2020-08-30 1561 citations 39
cs.LG 2008.12248

A Survey on Reinforcement Learning for Combinatorial Optimization

This survey reviews reinforcement learning (RL) methods for combinatorial optimization, focusing on TSP, from classical algorithms to deep RL with attention mechanisms, highlighting performance improvements.

Yunhao Yang, Andrew Whinston

2020-08-18 33 citations 34
cs.LG 2007.14062

Big Bird: Transformers for Longer Sequences

BigBird introduces sparse attention with global, local, and random tokens, reducing complexity from quadratic to linear, enabling long sequence modeling.

Manzil Zaheer, Guru Guruganesh, Avinava Dubey et al.

2020-07-28 49
cs.LG 2007.04612

Concept Bottleneck Models

Concept bottleneck models enable high-level concept prediction and intervention, maintaining competitive accuracy.

Pang Wei Koh, Thao Nguyen, Yew Siang Tang et al.

2020-07-09 46
cs.LG 2007.04309

Self-Supervised Policy Adaptation during Deployment

Self-supervised policy adaptation (PAD) enables reinforcement learning agents to online adapt in unseen environments without reward signals, significantly improving generalization.

Nicklas Hansen, Rishabh Jangir, Yu Sun et al.

2020-07-09 44