cs.CL 2302.01318

Accelerating Large Language Model Decoding with Speculative Sampling

Proposes Speculative Sampling, enabling 2-2.5x speedup in decoding 70B-parameter Chinchilla model by generating multiple tokens per call while maintaining distribution fidelity.

Charlie Chen, Sebastian Borgeaud, Geoffrey Irving et al.

2023-02-03 1069 citations 32
cs.CV 2302.00673

ADAPT: Action-aware Driving Caption Transformer

ADAPT, a Transformer-based multi-task model, jointly predicts driving actions and generates natural language explanations, achieving CIDEr 34.6 and reasoning 11.4 on BDD-X.

Bu Jin, Xinyu Liu, Yupeng Zheng et al.

2023-02-02 88
stat.ML 2301.13856

Simplex Random Features

Proposes Simplex Random Features (SimRFs) for optimal kernel approximation via geometric correlation, outperforming orthogonal RFs with minimal extra cost.

Isaac Reid, Krzysztof Choromanski, Valerii Likhosherstov et al.

2023-02-01 50
cs.LG 2301.13195

Adaptive Computation with Elastic Input Sequence

AdaTape enables dynamic computation in neural networks using adaptive tape tokens, excelling in image recognition tasks.

Fuzhao Xue, Valerii Likhosherstov, Anurag Arnab et al.

2023-01-31 9
cs.CL 2301.12652

REPLUG: Retrieval-Augmented Black-Box Language Models

REPLUG enhances black-box large models like GPT-3 by simple retrieval-based input augmentation, boosting performance by 6.3% without parameter fine-tuning.

Weijia Shi, Sewon Min, Michihiro Yasunaga et al.

2023-01-30 50
cs.LG 2301.11975

Byte Pair Encoding for Symbolic Music

Applying Byte Pair Encoding (BPE) to symbolic music reduces sequence length by over 50%, enlarges vocabulary, and improves model performance and speed.

Nathan Fradet, Nicolas Gutowski, Fabien Chhel et al.

2023-01-28 54