cs.LG 2407.17437

Nerva: a Truly Sparse Implementation of Neural Networks

Nerva leverages sparse matrix operations via Intel MKL to accelerate neural network training, reducing time by 4× at 99% sparsity with comparable accuracy.

Wieger Wesselink, Bram Grooten, Qiao Xiao et al.

2024-07-25 35
cs.LG 2407.12492

Temporal Test-Time Adaptation with State-Space Models

STAD employs Bayesian filtering within state-space models to adapt to gradually evolving temporal distribution shifts, outperforming traditional TTA methods.

Mona Schirmer, Dan Zhang, Eric Nalisnick

2024-07-17 5 citations 31
cs.LG 2407.11055

Knowledge boosting during low-latency inference

Proposes 'Knowledge Boosting' to enhance small models during low-latency inference using delayed hints from remote large models, achieving up to 3.53dB SI-SDR gain.

Vidya Srinivas, Malek Itani, Tuochao Chen et al.

2024-07-10 4 citations 54
cs.LG 2407.04622

On scalable oversight with weak LLMs judging strong LLMs

This paper introduces debate as a scalable oversight protocol, using weak LLMs to judge strong LLMs, showing superior performance in information-asymmetry tasks with up to 8% accuracy gain.

Zachary Kenton, Noah Y. Siegel, János Kramár et al.

2024-07-06 102 citations 41
cs.LG 2406.19384

The Remarkable Robustness of LLMs: Stages of Inference?

Layer swapping and deletion experiments reveal four universal inference stages in LLMs, showing high robustness especially in middle layers, with detailed neural mechanisms identified.

Vedang Lad, Jin Hwa Lee, Wes Gurnee et al.

2024-06-28 28
cs.LG 2406.19272

Stochastic Concept Bottleneck Models

Proposed Stochastic Concept Bottleneck Model (SCBM) models concept dependencies via multivariate normal distribution, significantly improving intervention effectiveness on synthetic and real datasets.

Moritz Vandenhirtz, Sonia Laguna, Ričards Marcinkevičs et al.

2024-06-27 36