Understanding deep learning requires rethinking generalization
Using random label/noise experiments, reveals deep networks' capacity to memorize, challenging classical generalization theories.
Chiyuan Zhang, Samy Bengio, Moritz Hardt et al.
Using random label/noise experiments, reveals deep networks' capacity to memorize, challenging classical generalization theories.
Chiyuan Zhang, Samy Bengio, Moritz Hardt et al.
Random-feature preconditioned PCG solves exact KRR, reaching high accuracy on up to one million examples in about one hour.
Haim Avron, Kenneth L. Clarkson, David P. Woodruff
Proposes Stein Sample Learning via Stein Variational Gradient Descent (SVGD), enabling neural networks to approximate target distributions without explicit density calculations, applied to deep energy models.
Dilin Wang, Qiang Liu
Differentiable physics engine enables gradient-based optimization of robot controllers, significantly improving training speed and scalability.
Jonas Degrave, Michiel Hermans, Joni Dambre et al.
Proposes Bidirectional Attention Flow (BIDAF), a multi-layer hierarchical model with bidirectional attention, significantly improving machine comprehension performance.
Minjoon Seo, Aniruddha Kembhavi, Ali Farhadi et al.
Evaluating LSTM's ability to learn syntax-sensitive dependencies via number prediction and language modeling experiments.
Tal Linzen, Emmanuel Dupoux, Yoav Goldberg
Proposes a scalable nonparametric Poisson intensity estimation method using transformed RKHS kernels, with theoretical guarantees and efficient computation.
Seth Flaxman, Yee Whye Teh, Dino Sejdinovic
End-to-end deep architecture based on R-MAC with triplet loss significantly improves image retrieval accuracy, achieving over 94% MAP on standard benchmarks.
Albert Gordo, Jon Almazan, Jerome Revaud et al.
3D-GAN generates high-quality 3D objects, enhancing recognition performance.
Jiajun Wu, Chengkai Zhang, Tianfan Xue et al.
Proposes C2ST, a classifier-based two-sample test with interpretability and strong statistical guarantees.
David Lopez-Paz, Maxime Oquab
Proposed a multi-month trained MRNN decoder with data augmentation, achieving significant robustness against neural variability, outperforming Kalman filters in non-human primates.
David Sussillo, Sergey D. Stavisky, Jonathan C. Kao et al.
Proposes a polynomial-time SDP relaxation for approximating Gromov-Hausdorff distance, acting as a pseudo-metric for large-scale shape comparison.
Soledad Villar, Afonso S. Bandeira, Andrew J. Blumberg et al.
Proposes a parameter-free Coin Betting-based meta algorithm (CBCE) with superior strongly adaptive regret bounds for changing environments.
Kwang-Sung Jun, Francesco Orabona, Rebecca Willett et al.
MMD-ResNet利用残差网络实现基于最大均值差异的非线性批次效应校准,有效减弱多源数据偏差。
Uri Shaham, Kelly P. Stanton, Jun Zhao et al.
Proposes Equation Learner (EQL), a neural network that learns analytical expressions and excels in extrapolation beyond training data.
Georg Martius, Christoph H. Lampert
Proposes Xception, a CNN based on depthwise separable convolutions, outperforming Inception V3 with similar parameters.
François Chollet
Introduced self-ensembling method, reducing SVHN error rate from 18.44% to 7.05%.
Samuli Laine, Timo Aila
Detect misclassified and out-of-distribution examples in neural networks using softmax probabilities, enhancing detection accuracy.
Dan Hendrycks, Kevin Gimpel
Proposes linear classifier probes to monitor intermediate features, revealing monotonic increase in linear separability across layers.
Guillaume Alain, Yoshua Bengio
Deep ReLU networks more efficiently approximate smooth functions in Sobolev spaces than shallow networks.
Dmitry Yarotsky