LaF: Labeling-Free Model Selection for Automated Deep Neural Network Reusing
LaF: Bayesian-based, labeling-free model ranking method, improves performance estimation without labels.
Qiang Hu, Yuejun Guo, Maxime Cordy et al.
LaF: Bayesian-based, labeling-free model ranking method, improves performance estimation without labels.
Qiang Hu, Yuejun Guo, Maxime Cordy et al.
This survey systematically reviews graph embedding methods, including traditional and GNN-based techniques for static and dynamic graphs, analyzing over 300 papers since 2017.
Shima Khoshraftar, Aijun An
t5x and seqio enable training of models with hundreds of billions of parameters on multi-terabyte datasets, using XLA GSPMD for efficient distributed parallelism.
Adam Roberts, Hyung Won Chung, Anselm Levskaya et al.
Introduces a doubly-robust estimator for position bias correction, reducing variance and improving unbiased ranking with fewer data.
Harrie Oosterhuis
Proposed a primal-dual kernelized bandit algorithm with sublinear regret and soft constraint violation guarantees, compatible with UCB, TS, and random exploration.
Xingyu Zhou, Bo Ji
SolidGen directly generates B-reps autoregressively, enabling conditional CAD synthesis without sequence supervision.
Pradeep Kumar Jayaraman, Joseph G. Lambourne, Nishkrit Desai et al.
CODEGEN uses multi-turn prompts to improve synthesis, reaching 47.34% on MTPB with its 16.1B model.
Erik Nijkamp, Bo Pang, Hiroaki Hayashi et al.
Proposes oscillation dampening and iterative weight freezing to improve low-bit quantization accuracy of models like MobileNetV2, achieving state-of-the-art results.
Markus Nagel, Marios Fournarakis, Yelysei Bondarenko et al.
This paper analyzes challenges in forecast evaluation, proposing best practices and pitfalls avoidance strategies.
Hansika Hewamalage, Klaus Ackermann, Christoph Bergmeir
Introduces a Hermite polynomial-based Gaussian Hermite DPP sampling method leveraging random matrix spectral density, improving Monte Carlo integration efficiency.
Nicholas P Baskerville
Model soups: averaging multiple fine-tuned models boosts accuracy to 90.94% on ImageNet without extra inference cost.
Mitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre et al.
GeoDiff employs a geometric diffusion model for direct atomic coordinate generation, outperforming SOTA especially on large molecules.
Minkai Xu, Lantao Yu, Yang Song et al.
Proposes safety and reliability weighted detection metrics using a target criticality model; evaluated nine detectors on nuScenes.
Andrea Ceccarelli, Leonardo Montecchi
Proposes a theory of abstraction with three key criteria to improve sample efficiency and generalization in reinforcement learning.
David Abel
This study reveals that fine-tuning can distort pretrained features, harming out-of-distribution generalization; proposes LP-FT strategy for balanced performance.
Ananya Kumar, Aditi Raghunathan, Robbie Jones et al.
Proposes the merged-staircase property (MSP) as a necessary and nearly sufficient condition for SGD learning of sparse functions on two-layer neural networks.
Emmanuel Abbe, Enric Boix-Adsera, Theodor Misiakiewicz
Proposes forward gradient method using forward automatic differentiation to estimate gradients without backpropagation, enabling training speeds up to twice as fast.
Atılım Güneş Baydin, Barak A. Pearlmutter, Don Syme et al.
This study quantifies memorization in large language models via log-linear relationships, showing that model size, data duplication, and context length significantly increase memorization risk.
Nicholas Carlini, Daphne Ippolito, Matthew Jagielski et al.
LPSDA leverages Lie point symmetries to enhance neural PDE solvers, reducing training data needs by over tenfold and improving generalization.
Johannes Brandstetter, Max Welling, Daniel E. Worrall
Recurrent neural networks with explicit memory and progressive training enable extreme algorithmic extrapolation, solving tasks beyond training scope with high accuracy.
Arpit Bansal, Avi Schwarzschild, Eitan Borgnia et al.