stat.CO 2203.15852

Direct Sampling with a Step Function

Proposes a step-function-based univariate direct sampling method that reduces manual tuning and rejection, enabling exact sampling from complex distributions.

Andrew M. Raim

2022-03-30 39
cs.LG 2203.15589

On Kernelized Multi-Armed Bandits with Constraints

Proposed a primal-dual kernelized bandit algorithm with sublinear regret and soft constraint violation guarantees, compatible with UCB, TS, and random exploration.

Xingyu Zhou, Bo Ji

2022-03-29 50
cs.CL 2203.15556

Training Compute-Optimal Large Language Models

This study introduces compute-optimal training for large language models, showing model size and data should scale together; trained 70B Chinchilla surpasses larger models.

Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch et al.

2022-03-29 24
cs.CV 2203.12119

Visual Prompt Tuning

This paper introduces Visual Prompt Tuning (VPT), which inserts less than 1% trainable prompts into the input space of frozen vision transformers, outperforming full fine-tuning on 20 out of 24 tasks with significantly fewer parameters.

Menglin Jia, Luming Tang, Bor-Chun Chen et al.

2022-03-23 2841 citations 41
cs.LG 2203.11086

Overcoming Oscillations in Quantization-Aware Training

Proposes oscillation dampening and iterative weight freezing to improve low-bit quantization accuracy of models like MobileNetV2, achieving state-of-the-art results.

Markus Nagel, Marios Fournarakis, Yelysei Bondarenko et al.

2022-03-22 45