Online Learning for Min Sum Set Cover and Pandora's Box
A convex-relaxation–FTRL–rounding framework achieves 9.22-approximate no-regret for online Pandora’s Box.
Evangelia Gergatsouli, Christos Tzamos
A convex-relaxation–FTRL–rounding framework achieves 9.22-approximate no-regret for online Pandora’s Box.
Evangelia Gergatsouli, Christos Tzamos
Proposes TACTiS, a Transformer-based model using attention copulas for joint probabilistic forecasting of high-dimensional multivariate time series.
Alexandre Drouin, Étienne Marcotte, Nicolas Chapados
GMC introduces a geometric contrastive framework with a two-level architecture, achieving state-of-the-art robustness in missing modality scenarios.
Petra Poklukar, Miguel Vasco, Hang Yin et al.
MP-PDE unifies classical local solvers with message passing and improves autoregressive stability via pushforward training.
Johannes Brandstetter, Daniel Worrall, Max Welling
Introduces a system of SDE-based score model for graph generation, capturing complex node-edge dependencies with high fidelity.
Jaehyeong Jo, Seul Lee, Sung Ju Hwang
Proposed ClippedGossip algorithm achieves neighborhood convergence for non-convex decentralized training under Byzantine attacks, leveraging spectral gap γ and attack ratio δ.
Lie He, Sai Praneeth Karimireddy, Martin Jaggi
Proposes Nash-MTL, modeling gradient aggregation as a bargaining game, achieving proportional fairness and improving multi-task learning performance.
Aviv Navon, Aviv Shamsian, Idan Achituve et al.
StyleGAN-XL uses Projected GAN strategy to successfully train on ImageNet, generating 1024² resolution images.
Axel Sauer, Katja Schwarz, Andreas Geiger
Proposes functa framework, representing data points as neural implicit functions for multi-modal deep learning.
Emilien Dupont, Hyunjik Kim, S. M. Ali Eslami et al.
Proposed GPFQ, a greedy path-following post-training quantization method with provable error bounds across diverse input distributions and architectures.
Jinjie Zhang, Yixuan Zhou, Rayan Saab
Introducing UOT-Pooling, a unified, learnable global pooling layer based on unbalanced optimal transport, improving performance across multiple tasks.
Minjie Cheng, Hongteng Xu
DeepSpeed-MoE advances MoE model efficiency with novel architecture and compression, achieving 4.5x faster and 9x cheaper inference.
Samyam Rajbhandari, Conglong Li, Zhewei Yao et al.
Studied 'grokking' phenomenon in small algorithm datasets; regularization (like weight decay) accelerates sudden generalization after overfitting.
Alethea Power, Yuri Burda, Harri Edwards et al.
Proposes ECOLog, an efficient and statistically optimal Logistic Bandit algorithm, achieving O(d^2) per-round complexity with regret matching the lower bound.
Louis Faury, Marc Abeille, Kwang-Sung Jun et al.
Proposes FedLRGD leveraging data smoothness for lower federated oracle complexity than FedAve.
Ali Jadbabaie, Anuran Makur, Devavrat Shah
L2P introduces a prompt pool with instance-wise query for continual learning without task labels, outperforming buffer-based methods.
Zifeng Wang, Zizhao Zhang, Chen-Yu Lee et al.
Proposes a PIT entropy-based recalibration method for epidemic forecasts, significantly improving calibration and accuracy across 27 influenza models.
Aaron Rumack, Ryan J. Tibshirani, Roni Rosenfeld
Proposes O(1) memory attention algorithm extended to O(log n), enabling scalable long-sequence modeling.
Markus N. Rabe, Charles Staats
Proposes persistent homology dimension (PHD) as a topological measure to estimate neural network intrinsic dimension and predict generalization error.
Tolga Birdal, Aaron Lou, Leonidas Guibas et al.
Scatterbrain unifies sparse and low-rank attention approximation using LSH and kernel features, reducing error by 2.1× over baselines.
Beidi Chen, Tri Dao, Eric Winsor et al.