Entire Space Multi-Task Model: An Effective Approach for Estimating Post-Click Conversion Rate
Proposed the Entire Space Multi-Task Model (ESMM), achieving a 2.56% AUC improvement on Taobao dataset.
Xiao Ma, Liqin Zhao, Guan Huang et al.
Proposed the Entire Space Multi-Task Model (ESMM), achieving a 2.56% AUC improvement on Taobao dataset.
Xiao Ma, Liqin Zhao, Guan Huang et al.
ODIN integrates 64k synapses and 256 neurons in 28nm CMOS, achieving 12.7pJ/SOP with online learning and multi-behavior models.
Charlotte Frenkel, Martin Lefebvre, Jean-Didier Legat et al.
Proposed partial convolution for image inpainting, achieving 33.75 PSNR and reducing artifacts in irregular masks.
Guilin Liu, Fitsum A. Reda, Kevin J. Shih et al.
GLUE benchmark with 9 tasks, multi-task training improves generalization; baseline scores still far from human performance.
Alex Wang, Amanpreet Singh, Julian Michael et al.
Introduces Simulation-Based Calibration (SBC) for validating Bayesian inference algorithms via posterior sample self-consistency, detecting biases and implementation errors.
Sean Talts, Michael Betancourt, Daniel Simpson et al.
Proposes distributional dynamics (DD) as a PDE framework to analyze SGD in two-layer neural networks, proving convergence in large-scale limits.
Song Mei, Andrea Montanari, Phan-Minh Nguyen
Pre-trained word embeddings improve low-resource NMT by up to 20 BLEU points.
Ye Qi, Devendra Singh Sachan, Matthieu Felix et al.
Proposes Dual Learning Algorithm (DLA) for joint unbiased estimation of propensity and ranking models, outperforming randomized unbiased methods.
Qingyao Ai, Keping Bi, Cheng Luo et al.
Proposed a discourse-aware model for long document summarization, significantly outperforming existing models.
Arman Cohan, Franck Dernoncourt, Doo Soon Kim et al.
Enhanced quantum SDP solver using improved Gibbs samplers, applied to shadow tomography.
Joran van Apeldoorn, András Gilyén
This paper analyzes the impact of hyperparameters in Word2Vec's Skip-gram negative sampling for recommendation, showing significant performance gains through optimization.
Hugo Caselles-Dupré, Florian Lesaint, Jimena Royo-Letelier
Online distillation enables parallel training of multiple models, doubling training speed and improving reproducibility on large datasets.
Rohan Anil, Gabriel Pereyra, Alexandre Passos et al.
Introduces Speech Commands dataset with 105,829 samples for keyword spotting; achieves up to 88.2% accuracy with CNN models.
Pete Warden
Pixel2Mesh employs graph convolutional networks to deform an initial ellipsoid into detailed 3D meshes from a single RGB image, outperforming state-of-the-art by significant margins.
Nanyang Wang, Yinda Zhang, Zhuwen Li et al.
Using internet multi-view images with SfM and MVS to create MegaDepth, greatly enhancing generalization in single-view depth prediction.
Zhengqi Li, Noah Snavely
Jacquard dataset uses simulated environments to generate large-scale grasp locations, enhancing robotic grasp detection performance.
Amaury Depierre, Emmanuel Dellandréa, Liming Chen
GrBAL and ReBAL enable rapid online adaptation with only 1.5–3 hours of meta-training experience.
Anusha Nagabandi, Ignasi Clavera, Simin Liu et al.
This paper demonstrates that random parameterized quantum circuits exhibit exponential gradient vanishing (barren plateaus) in high-dimensional Hilbert space, severely limiting the effectiveness of gradient-based training.
Jarrod R. McClean, Sergio Boixo, Vadim N. Smelyanskiy et al.
Proposes Point Convolutional Neural Networks (PCNN) using extension-restriction operators, achieving invariant, robust point cloud convolution with state-of-the-art results.
Matan Atzmon, Haggai Maron, Yaron Lipman
Using neural network-based word embeddings to analyze cultural dimensions via geometric relationships in high-dimensional space.
Austin C. Kozlowski, Matt Taddy, James A. Evans