cs.LG 2301.07733

Learning-Rate-Free Learning by D-Adaptation

D-Adaptation achieves parameter-free optimal convergence via adaptive D-estimation, matching hand-tuned rates without line searches.

Aaron Defazio, Konstantin Mishchenko

2023-01-19 25
cs.LG 2301.05860

State of the Art and Potentialities of Graph-level Learning

This survey presents a comprehensive taxonomy of graph-level learning methods, including traditional kernels, substructure mining, graph embeddings, GNNs, and pooling, with performance insights.

Zhenyu Yang, Ge Zhang, Jia Wu et al.

2023-01-14 21
cs.LG 2212.09720

The case for 4-bit precision: k-bit Inference Scaling Laws

This study establishes that 4-bit quantization offers near-universal optimality for zero-shot performance and model size trade-offs in large language models, validated by 35,000 experiments.

Tim Dettmers, Luke Zettlemoyer

2022-12-20 34
cs.LG 2211.05244

Deep Learning for Time Series Anomaly Detection: A Survey

Proposes a comprehensive taxonomy of deep learning models for time series anomaly detection, covering forecasting, reconstruction, representation, and hybrid methods.

Zahra Zamanzadeh Darban, Geoffrey I. Webb, Shirui Pan et al.

2022-11-10 30
cs.LG 2211.05102

Efficiently Scaling Transformer Inference

Proposes multi-dimensional partitioning and communication optimization for TPU v4, achieving 29ms/token latency and 76% MFU on 540B models.

Reiner Pope, Sholto Douglas, Aakanksha Chowdhery et al.

2022-11-10 43
cs.LG 2210.11416

Scaling Instruction-Finetuned Language Models

Scaling instruction fine-tuning with 1.8K tasks and chain-of-thought enhances model performance, achieving SOTA on multiple benchmarks.

Hyung Won Chung, Le Hou, Shayne Longpre et al.

2022-10-21 54
cs.LG 2210.10760

Scaling Laws for Reward Model Overoptimization

Using synthetic setup, the study models reward overoptimization laws; relationships differ between RL and BoN, coefficients scale smoothly with model size.

Leo Gao, John Schulman, Jacob Hilton

2022-10-20 27