cs.LG 2406.11235

QTIP: Quantization with Trellises and Incoherence Processing

QTIP employs trellis-coded quantization (TCQ) with hardware-efficient codes to enable ultra-high-dimensional weight quantization, boosting model compression and inference speed.

Albert Tseng, Qingyao Sun, David Hou et al.

2024-06-17 41
cs.LG 2406.07887

An Empirical Study of Mamba-based Language Models

This study compares 8B Mamba, Mamba-2, and Transformer models trained on up to 3.5T tokens, showing hybrid models outperform in speed and long-sequence tasks.

Roger Waleffe, Wonmin Byeon, Duncan Riach et al.

2024-06-12 35
cs.LG 2406.04843

Variational Flow Matching for Graph Generation

Proposes Variational Flow Matching (VFM) and CatFlow for categorical graph generation, outperforming current SOTA.

Floor Eijkelboom, Grigory Bartosh, Christian Andersson Naesseth et al.

2024-06-07 33
cs.LG 2406.04713

FlowMM: Generating Materials with Riemannian Flow Matching

FlowMM employs Riemannian Flow Matching to generate stable crystal structures efficiently, reducing inference steps by 3x compared to diffusion models.

Benjamin Kurt Miller, Ricky T. Q. Chen, Anuroop Sriram et al.

2024-06-07 36
cs.LG 2406.04446

Can Language Models Use Forecasting Strategies?

This study evaluates PaLM 2's forecasting ability using a real-world event dataset, revealing a bias toward underestimating event probabilities.

Sarah Pratt, Seth Blumberg, Pietro Kreitlon Carolino et al.

2024-06-07 16 citations 38
cs.LG 2406.04093

Scaling and evaluating sparse autoencoders

Using k-sparse autoencoders with TopK activation, trained on GPT-4 activations, scaled to 16 million latent variables, revealing clear power-law relationships between size and feature quality.

Leo Gao, Tom Dupré la Tour, Henk Tillman et al.

2024-06-06 567 citations 42