Minimum Risk Training for Neural Machine Translation
Proposes Minimum Risk Training (MRT) for end-to-end neural machine translation, outperforming maximum likelihood estimation with up to +8.61 BLEU points.
Shiqi Shen, Yong Cheng, Zhongjun He et al.
Proposes Minimum Risk Training (MRT) for end-to-end neural machine translation, outperforming maximum likelihood estimation with up to +8.61 BLEU points.
Shiqi Shen, Yong Cheng, Zhongjun He et al.
Convolutional Neural Networks (CNNs) use convolutional and pooling layers for efficient image recognition.
Keiron O'Shea, Ryan Nash
Neural GPU uses convolutional GRUs for efficient algorithm learning, handling long inputs.
Łukasz Kaiser, Ilya Sutskever
Proposes PPDB-based universal paraphrastic sentence embeddings; simple models outperform LSTMs in cross-domain tasks.
John Wieting, Mohit Bansal, Kevin Gimpel et al.
Proposes NetVLAD, an end-to-end trainable CNN with a differentiable VLAD layer, achieving significant improvements in large-scale place recognition.
Relja Arandjelović, Petr Gronat, Akihiko Torii et al.
Multi-scale context aggregation using dilated convolutions improves semantic segmentation accuracy.
Fisher Yu, Vladlen Koltun
Recurrent neural network-based session recommender, outperforming traditional methods with 20%+ improvements in Recall@20.
Balázs Hidasi, Alexandros Karatzoglou, Linas Baltrunas et al.
Introduced Bilateral Inception module for superpixel CNNs, enhancing semantic segmentation accuracy.
Raghudeep Gadde, Varun Jampani, Martin Kiefel et al.
Neural Programmer-Interpreter (NPI) combines LSTM core, persistent program memory, and environment encoders to enable multi-task program learning and generalization.
Scott Reed, Nando de Freitas
OpenMax layer combined with Meta-Recognition estimates unknown class probability, improving open set recognition and rejecting fooling/adversarial samples.
Abhijit Bendale, Terrance Boult
Prioritized Experience Replay enhances DQN learning efficiency, outperforming baseline in 41 out of 49 games.
Tom Schaul, John Quan, Ioannis Antonoglou et al.
Frank-Wolfe variants achieve global linear convergence, effective for flow polytope constraints.
Simon Lacoste-Julien, Martin Jaggi
Gated Graph Sequence Neural Networks use GRUs and modern optimization to enhance sequence output from graph-structured data.
Yujia Li, Daniel Tarlow, Marc Brockschmidt et al.
Proposes Structural-RNN (S-RNN), combining high-level spatio-temporal graphs with RNNs, improving modeling of human motion and object interactions.
Ashesh Jain, Amir R. Zamir, Silvio Savarese et al.
Proposal Flow leverages multi-scale object proposals and geometric constraints for image correspondence, outperforming existing semantic flow methods.
Bumsub Ham, Minsu Cho, Cordelia Schmid et al.
Defensive distillation cuts adversarial success from 95.89% to 0.45%.
Nicolas Papernot, Patrick McDaniel, Xi Wu et al.
Introduces a polynomial-time algorithm for optimizing star-convex functions, overcoming gradient dependence and smoothness constraints.
Jasper C. H. Lee, Paul Valiant
Geometric analysis reveals that wider neural networks have higher probability of initializing in basins with low objective values, facilitating optimization.
Itay Safran, Ohad Shamir
Semantic segmentation using CNN and Domain Transform, achieving 66.35% mIOU with improved boundary precision.
Liang-Chieh Chen, Jonathan T. Barron, George Papandreou et al.
Neural Module Networks combine deep learning with linguistic structure, achieving top results on VQA datasets.
Jacob Andreas, Marcus Rohrbach, Trevor Darrell et al.