Rethinking the Expressive Power of GNNs via Graph Biconnectivity
Proposes GD-WL framework incorporating distance info, achieving provable expressiveness for graph biconnectivity tasks.
Bohang Zhang, Shengjie Luo, Liwei Wang et al.
Proposes GD-WL framework incorporating distance info, achieving provable expressiveness for graph biconnectivity tasks.
Bohang Zhang, Shengjie Luo, Liwei Wang et al.
D-Adaptation achieves parameter-free optimal convergence via adaptive D-estimation, matching hand-tuned rates without line searches.
Aaron Defazio, Konstantin Mishchenko
Symbolic expression generation via VAE, SEGVAE achieves 65% recovery rate on Nguyen dataset.
Sergei Popov, Mikhail Lazarev, Vladislav Belavin et al.
This survey presents a comprehensive taxonomy of graph-level learning methods, including traditional kernels, substructure mining, graph embeddings, GNNs, and pooling, with performance insights.
Zhenyu Yang, Ge Zhang, Jia Wu et al.
Mechanistic interpretability reveals continuous progress measures for grokking via Fourier transforms, showing gradual amplification of learned algorithms.
Neel Nanda, Lawrence Chan, Tom Lieberum et al.
AER model combines auto-encoder and regression, achieving a 23.5% F1 score improvement for time series anomaly detection.
Lawrence Wong, Dongyu Liu, Laure Berti-Equille et al.
This study establishes that 4-bit quantization offers near-universal optimality for zero-shot performance and model size trade-offs in large language models, validated by 35,000 experiments.
Tim Dettmers, Luke Zettlemoyer
Introduced continuous-time discrete diffusion models, enhancing music and image data generation.
Haoran Sun, Lijun Yu, Bo Dai et al.
Applying conditional diffusion models (Decision Diffuser) for offline decision-making surpasses traditional RL, enabling direct trajectory sampling with multi-condition support.
Anurag Ajay, Yilun Du, Abhi Gupta et al.
Proposes GO-UCB, a parametric model-based method achieving \~O(√T) regret for high-dimensional global optimization.
Chong Liu, Yu-Xiang Wang
Proposes a comprehensive taxonomy of deep learning models for time series anomaly detection, covering forecasting, reconstruction, representation, and hybrid methods.
Zahra Zamanzadeh Darban, Geoffrey I. Webb, Shirui Pan et al.
Proposes multi-dimensional partitioning and communication optimization for TPU v4, achieving 29ms/token latency and 76% MFU on 540B models.
Reiner Pope, Sholto Douglas, Aakanksha Chowdhery et al.
DPM-Solver++: Data-prediction-based high-order solver, generates high-quality images in 15 steps with guided diffusion models.
Cheng Lu, Yuhao Zhou, Fan Bao et al.
GPTQ is a second-order based post-training quantization method that compresses 175B parameter models to 4 bits with minimal accuracy loss.
Elias Frantar, Saleh Ashkboos, Torsten Hoefler et al.
Survey of deep neural network methods (PINNs, Neural Operators) for solving PDEs, highlighting algorithms, applications, and future directions.
Shudong Huang, Wentao Feng, Chenwei Tang et al.
Algorithm Distillation (AD) uses causal sequence models to convert RL algorithms into neural networks, enhancing data efficiency.
Michael Laskin, Luyu Wang, Junhyuk Oh et al.
Using a GPT variant trained on synthetic Othello sequences, the study uncovers emergent nonlinear internal representations of the board state, validated through intervention experiments.
Kenneth Li, Aspen K. Hopkins, David Bau et al.
Scaling instruction fine-tuning with 1.8K tasks and chain-of-thought enhances model performance, achieving SOTA on multiple benchmarks.
Hyung Won Chung, Le Hou, Shayne Longpre et al.
Using synthetic setup, the study models reward overoptimization laws; relationships differ between RL and BoN, coefficients scale smoothly with model size.
Leo Gao, John Schulman, Jacob Hilton
ProtoVAE is a variational autoencoder-based self-explainable prototype model achieving high accuracy and trustworthy explanations.
Srishti Gautam, Ahcene Boubekki, Stine Hansen et al.