Understanding Neural Networks Through Deep Visualization
Deep visualization tools reveal computations in intermediate layers of convolutional neural networks.
Jason Yosinski, Jeff Clune, Anh Nguyen et al.
Deep visualization tools reveal computations in intermediate layers of convolutional neural networks.
Jason Yosinski, Jeff Clune, Anh Nguyen et al.
Proposes a graph-based visual relationship model using large-scale Amazon data for clothing and accessory recommendations.
Julian McAuley, Christopher Targett, Qinfeng Shi et al.
Unsupervised visual representation learning via context prediction, enhancing object discovery on Pascal VOC 2011 dataset.
Carl Doersch, Abhinav Gupta, Alexei A. Efros
Proposes CNN-based joint instance segmentation and depth ordering via multi-scale patch prediction and MRF fusion, achieving state-of-the-art on KITTI.
Ziyu Zhang, Alexander G. Schwing, Sanja Fidler et al.
Proposes a multi-region CNN with semantic segmentation-aware features and iterative bounding box refinement, achieving 78.2% mAP on VOC2007.
Spyros Gidaris, Nikos Komodakis
Neural-Image-QA combines CNN and LSTM, doubling previous accuracy to 17.49%, advancing multi-modal visual question answering.
Mateusz Malinowski, Marcus Rohrbach, Mario Fritz
Proposes a deep convolutional neural network-based direct perception model estimating 13 key affordance indicators for autonomous driving, trained on 12 hours of video game data.
Chenyi Chen, Ari Seff, Alain Kornhauser et al.
Fast R-CNN achieves 9× faster training, 146× faster detection, with 66% mAP on VOC2007, by sharing features and end-to-end multi-task learning.
Ross Girshick
FlowNet learns dense optical flow end-to-end, training on Flying Chairs and reaching 5–10 fps while generalizing to Sintel and KITTI.
Philipp Fischer, Alexey Dosovitskiy, Eddy Ilg et al.
Introduced HED, a new edge detection algorithm achieving an ODS F-score of 0.782 on the BSD500 dataset.
Saining Xie, Zhuowen Tu
DC-IGN learns interpretable image representations using SGVB, generating images with varied poses and lighting.
Tejas D. Kulkarni, Will Whitney, Pushmeet Kohli et al.
Proposed Deep Convolutional Neural Field model outperforms existing methods in monocular depth estimation.
Fayao Liu, Chunhua Shen, Guosheng Lin et al.
MatConvNet is a MATLAB-based CNN toolbox supporting fast prototyping and GPU acceleration, enabling training on large datasets like ImageNet.
Andrea Vedaldi, Karel Lenc
Hypercolumns combine multi-layer CNN features for pixel-level object segmentation, boosting SDS from 49.7 to 60.0 mean APr.
Bharath Hariharan, Pablo Arbeláez, Ross Girshick et al.
Proposes a deep neural image captioning model combining CNN and LSTM, achieving BLEU-4 of 27.7, outperforming previous methods.
Oriol Vinyals, Alexander Toshev, Samy Bengio et al.
Proposes LRCN, combining CNN and LSTM for video recognition and captioning, outperforming single models with end-to-end training.
Jeff Donahue, Lisa Anne Hendricks, Marcus Rohrbach et al.
Proposes Fully Convolutional Networks (FCN) for pixel-level semantic segmentation, achieving 62.2% mean IU, surpassing previous SOTA with end-to-end training.
Jonathan Long, Evan Shelhamer, Trevor Darrell
Dictionary learning combined with `1-norm sparse coding improves image representation and recognition accuracy.
Julien Mairal, Francis Bach, Jean Ponce
Proposed Inception architecture employs multi-scale convolutions and dimension reduction, achieving 28.6% Top-5 error on ImageNet with 1/12 parameters of AlexNet.
Christian Szegedy, Wei Liu, Yangqing Jia et al.
Deep CNNs like AlexNet, VGG, ResNet trained on 14 million images achieved top-5 error rates below 7%, revolutionizing large-scale image recognition.
Olga Russakovsky, Jia Deng, Hao Su et al.