stat.ME 2106.06137

Conformal Bayesian Computation

Proposes scalable conformal Bayesian prediction intervals using importance sampling, ensuring finite-sample coverage with reduced computational cost.

Edwin Fong, Chris Holmes

2021-06-11 37
cs.CV 2106.05974

Scaling Vision with Sparse Mixture of Experts

Introduces sparse MoE into Vision Transformer, achieving 90.35% on ImageNet with 15B parameters, halving inference compute compared to dense models.

Carlos Riquelme, Joan Puigcerver, Basil Mustafa et al.

2021-06-11 37
cs.CV 2106.04531

RobustNav: Towards Benchmarking Robustness in Embodied Navigation

RobustNav benchmarks embodied navigation robustness under visual and dynamics corruptions, revealing significant performance drops and highlighting the need for improved adaptation.

Prithvijit Chattopadhyay, Judy Hoffman, Roozbeh Mottaghi et al.

2021-06-09 38
cs.LG 2106.04426

Hash Layers For Large Sparse Models

Efficient parameter allocation in large sparse models using Hash Layers, improving performance.

Stephen Roller, Sainbayar Sukhbaatar, Arthur Szlam et al.

2021-06-08 10
cs.LG 2106.03893

Rethinking Graph Transformers with Spectral Attention

Spectral Attention Network (SAN) leverages Laplacian spectrum for node positional encoding, outperforming traditional GNNs with full connectivity.

Devin Kreuzer, Dominique Beaini, William L. Hamilton et al.

2021-06-08 38
cs.LG 2106.03764

On the Expressive Power of Self-Attention Matrices

This paper proves self-attention matrices can approximate arbitrary sparse patterns with input adjustment, requiring hidden size d = O(log L).

Valerii Likhosherstov, Krzysztof Choromanski, Adrian Weller

2021-06-08 45