Unsupervised Visual Representation Learning by Context Prediction
Unsupervised visual representation learning via context prediction, enhancing object discovery on Pascal VOC 2011 dataset.
Carl Doersch, Abhinav Gupta, Alexei A. Efros
Unsupervised visual representation learning via context prediction, enhancing object discovery on Pascal VOC 2011 dataset.
Carl Doersch, Abhinav Gupta, Alexei A. Efros
Proposes CNN-based joint instance segmentation and depth ordering via multi-scale patch prediction and MRF fusion, achieving state-of-the-art on KITTI.
Ziyu Zhang, Alexander G. Schwing, Sanja Fidler et al.
Proposes a multi-region CNN with semantic segmentation-aware features and iterative bounding box refinement, achieving 78.2% mAP on VOC2007.
Spyros Gidaris, Nikos Komodakis
Neural-Image-QA combines CNN and LSTM, doubling previous accuracy to 17.49%, advancing multi-modal visual question answering.
Mateusz Malinowski, Marcus Rohrbach, Mario Fritz
Highway Networks enable seamless information flow via gating units, supporting training of hundreds of layers deep networks.
Rupesh Kumar Srivastava, Klaus Greff, Jürgen Schmidhuber
Proposes a deep convolutional neural network-based direct perception model estimating 13 key affordance indicators for autonomous driving, trained on 12 hours of video game data.
Chenyi Chen, Ari Seff, Alain Kornhauser et al.
Fast R-CNN achieves 9× faster training, 146× faster detection, with 66% mAP on VOC2007, by sharing features and end-to-end multi-task learning.
Ross Girshick
FlowNet learns dense optical flow end-to-end, training on Flying Chairs and reaching 5–10 fps while generalizing to Sintel and KITTI.
Philipp Fischer, Alexey Dosovitskiy, Eddy Ilg et al.
Introduced HED, a new edge detection algorithm achieving an ODS F-score of 0.782 on the BSD500 dataset.
Saining Xie, Zhuowen Tu
MAP-Elites maps high-performance solutions across feature space, revealing solution distribution and diversity.
Jean-Baptiste Mouret, Jeff Clune
Proposes incremental sparse Gaussian process regression for continuous-time trajectory estimation, achieving 3x speedup while maintaining accuracy.
Xinyan Yan, Vadim Indelman, Byron Boots
Kernel Manifold Alignment (KEMA) enables multi-source, unpaired domain alignment with superior performance on synthetic and real datasets.
Devis Tuia, Gustau Camps-Valls
Proposes end-to-end training of deep visuomotor policies using Guided Policy Search with a 92,000-parameter CNN for direct image-to-torque mapping.
Sergey Levine, Chelsea Finn, Trevor Darrell et al.
K-FAC approximates Fisher matrix via Kronecker decomposition, enabling faster natural gradient optimization in neural networks.
James Martens, Roger Grosse
LINE efficiently embeds large-scale networks by optimizing first- and second-order proximities with edge sampling, handling millions of nodes and billions of edges.
Jian Tang, Meng Qu, Mingzhe Wang et al.
DC-IGN learns interpretable image representations using SGVB, generating images with varied poses and lighting.
Tejas D. Kulkarni, Will Whitney, Pushmeet Kohli et al.
Knowledge distillation transfers ensemble model knowledge into a single small model, achieving near-ensemble performance on MNIST and speech recognition tasks.
Geoffrey Hinton, Oriol Vinyals, Jeff Dean
The Bayesian Case Model (BCM) integrates case-based reasoning with a generative framework.
Been Kim, Cynthia Rudin, Julie Shah
Support function-based online convex optimization achieves Blackwell approachability with O(T−1/2) convergence.
Nahum Shimkin
Exp3.G classifies feedback graphs: strongly observable gives ~√(αT), weakly observable ~δ^(1/3)T^(2/3), and unobservable Θ(T).
Noga Alon, Nicolò Cesa-Bianchi, Ofer Dekel et al.