Entire Space Multi-Task Model: An Effective Approach for Estimating Post-Click Conversion Rate
Proposed the Entire Space Multi-Task Model (ESMM), achieving a 2.56% AUC improvement on Taobao dataset.
Xiao Ma, Liqin Zhao, Guan Huang et al.
Proposed the Entire Space Multi-Task Model (ESMM), achieving a 2.56% AUC improvement on Taobao dataset.
Xiao Ma, Liqin Zhao, Guan Huang et al.
Proposes distributional dynamics (DD) as a PDE framework to analyze SGD in two-layer neural networks, proving convergence in large-scale limits.
Song Mei, Andrea Montanari, Phan-Minh Nguyen
Transformer-based attention model trained with REINFORCE and greedy rollout baseline, achieving near-optimal solutions for TSP and VRP with node counts up to 100, outperforming previous learned heuristics.
Wouter Kool, Herke van Hoof, Max Welling
Learn unknown ODE models using Gaussian processes to infer dynamics from sparse data and predict future states.
Markus Heinonen, Cagatay Yildiz, Henrik Mannerström et al.
Proposes NO TEARS, a continuous optimization framework for DAG structure learning using matrix exponential-based smooth constraints, outperforming traditional combinatorial methods.
Xun Zheng, Bryon Aragam, Pradeep Ravikumar et al.
Efficient numerical methods for optimal transport enable scalable high-dimensional distribution matching, benefiting image processing and machine learning.
Gabriel Peyré, Marco Cuturi
Proposed Fast Geometric Ensembling (FGE) improves accuracy by 0.56% on CIFAR-10.
Timur Garipov, Pavel Izmailov, Dmitrii Podoprikhin et al.
Using input-output Jacobian norm as a sensitivity metric, the study finds strong correlation with neural network generalization across architectures and datasets.
Roman Novak, Yasaman Bahri, Daniel A. Abolafia et al.
Introduces Oblivious-Greedy for non-submodular maximization under element removal, achieving constant-factor approximation for support selection and variance reduction.
Ilija Bogunovic, Junyao Zhao, Volkan Cevher
FactorVAE improves disentanglement over β-VAE by encouraging factorial representation distribution.
Hyunjik Kim, Andriy Mnih
UMAP leverages Riemannian geometry and topology to produce scalable, high-quality low-dimensional embeddings that preserve both local and global data structures.
Leland McInnes, John Healy, James Melville
This paper clarifies bias issues in MMD-GAN training, introduces Kernel Inception Distance for evaluation, and demonstrates improved stability and efficiency.
Mikołaj Bińkowski, Danica J. Sutherland, Michael Arbel et al.
Proposes a noise-tolerant loss function based on Mean Absolute Error (MAE), theoretically proven to enable risk minimization to learn true classifiers under label noise in multi-class classification.
Aritra Ghosh, Himanshu Kumar, P. S. Sastry
Introduces Concept Activation Vectors (CAV) and TCAV to quantify high-level concept importance in neural networks, validated on image and medical datasets, demonstrating interpretability and bias detection.
Been Kim, Martin Wattenberg, Justin Gilmer et al.
Introduces bi-stochastic kernels for diffusion processes on manifolds, deriving their infinitesimal generators and heat kernel connections, with spectral and Nyström analysis.
Nicholas F. Marshall, Ronald R. Coifman
Proposes a two-step approach: stochastic dual OT plan learning and neural network Monge map approximation, applied to domain adaptation and generative modeling.
Vivien Seguy, Bharath Bhushan Damodaran, Rémi Flamary et al.
Variational Walkback learns non-equilibrium transition operators to generate samples, excelling on datasets like MNIST.
Anirudh Goyal, Nan Rosemary Ke, Surya Ganguli et al.
Proposes SNPE, a neural network-based ABC method for full Bayesian inference of neural models, accurately recovering parameters and handling missing data.
Jan-Matthis Lueckmann, Pedro J. Goncalves, Giacomo Bassetto et al.
Proves gradient descent on separable data converges directionally to max-margin (hard margin SVM) solution with rate O(1/ log t).
Daniel Soudry, Elad Hoffer, Mor Shpigel Nacson et al.
Proposed a new estimator combining logistic regression and Kullback-Leibler distance to efficiently estimate conversion probability on Criteo dataset.
Abdollah Safari, Rachel MacKay Altman, Thomas M. Loughin