cs.LG 2106.03893

Rethinking Graph Transformers with Spectral Attention

Spectral Attention Network (SAN) leverages Laplacian spectrum for node positional encoding, outperforming traditional GNNs with full connectivity.

Devin Kreuzer, Dominique Beaini, William L. Hamilton et al.

2021-06-08 30
cs.LG 2106.03764

On the Expressive Power of Self-Attention Matrices

This paper proves self-attention matrices can approximate arbitrary sparse patterns with input adjustment, requiring hidden size d = O(log L).

Valerii Likhosherstov, Krzysztof Choromanski, Adrian Weller

2021-06-08 39
cs.LG 2106.02711

SketchGen: Generating Constrained CAD Sketches

SketchGen uses Transformer-based sequence modeling to generate CAD sketches, outperforming state-of-the-art methods with improved distribution alignment.

Wamiq Reyaz Para, Shariq Farooq Bhat, Paul Guerrero et al.

2021-06-05 29
cs.LG 2105.14995

Choose a Transformer: Fourier or Galerkin

Proposes Galerkin Transformer with softmax-free attention, grounded in Petrov-Galerkin theory, improving PDE operator learning efficiency.

Shuhao Cao

2021-05-31 46
cs.LG 2105.13345

Adversarial Intrinsic Motivation for Reinforcement Learning

AIM leverages Wasserstein-1 distance with a goal-specific quasimetric to efficiently guide goal-conditioned reinforcement learning, accelerating convergence by ~50%.

Ishan Durugkar, Mauricio Tec, Scott Niekum et al.

2021-05-28 34
cs.LG 2105.06331

Informed Equation Learning

Informed Equation Learning (iEQL) integrates expert knowledge and robust training strategies to support singular functions like log and division, achieving high-accuracy, interpretable models.

Matthias Werner, Andrej Junginger, Philipp Hennig et al.

2021-05-13 21 citations 33
cs.LG 2105.05233

Diffusion Models Beat GANs on Image Synthesis

Diffusion models outperform GANs in image synthesis, achieving FID scores of 2.97, 4.59, and 7.72 on ImageNet 128×128, 256×256, and 512×512 respectively, with efficient sampling and classifier guidance.

Prafulla Dhariwal, Alex Nichol

2021-05-12 12901 citations 47
cs.LG 2104.13906

Reward (Mis)design for Autonomous Driving

Proposes 8 sanity checks for reward functions, revealing widespread flaws in autonomous driving RL reward design.

W. Bradley Knox, Alessandro Allievi, Holger Banzhaf et al.

2021-04-29 176 citations 45
cs.LG 2104.10350

Carbon Emissions and Large Neural Network Training

This study quantifies the energy use and carbon footprint of large NLP models like GPT-3, proposing strategies for greener training.

David Patterson, Joseph Gonzalez, Quoc Le et al.

2021-04-21 28