cs.CV 1504.08083

Fast R-CNN

Fast R-CNN achieves 9× faster training, 146× faster detection, with 66% mAP on VOC2007, by sharing features and end-to-end multi-task learning.

Ross Girshick

2015-04-30 51
cs.CV 1504.06375

Holistically-Nested Edge Detection

Introduced HED, a new edge detection algorithm achieving an ODS F-score of 0.782 on the BSD500 dataset.

Saining Xie, Zhuowen Tu

2015-04-24 43
cs.CV 1503.03167

Deep Convolutional Inverse Graphics Network

DC-IGN learns interpretable image representations using SGVB, generating images with varied poses and lighting.

Tejas D. Kulkarni, Will Whitney, Pushmeet Kohli et al.

2015-03-11 4
cs.CV 1411.4555

Show and Tell: A Neural Image Caption Generator

Proposes a deep neural image captioning model combining CNN and LSTM, achieving BLEU-4 of 27.7, outperforming previous methods.

Oriol Vinyals, Alexander Toshev, Samy Bengio et al.

2014-11-18 79
cs.CV 1411.4038

Fully Convolutional Networks for Semantic Segmentation

Proposes Fully Convolutional Networks (FCN) for pixel-level semantic segmentation, achieving 62.2% mean IU, surpassing previous SOTA with end-to-end training.

Jonathan Long, Evan Shelhamer, Trevor Darrell

2014-11-15 54
cs.CV 1409.4842

Going Deeper with Convolutions

Proposed Inception architecture employs multi-scale convolutions and dimension reduction, achieving 28.6% Top-5 error on ImageNet with 1/12 parameters of AlexNet.

Christian Szegedy, Wei Liu, Yangqing Jia et al.

2014-09-17 44
cs.CV 1409.0575

ImageNet Large Scale Visual Recognition Challenge

Deep CNNs like AlexNet, VGG, ResNet trained on 14 million images achieved top-5 error rates below 7%, revolutionizing large-scale image recognition.

Olga Russakovsky, Jia Deng, Hao Su et al.

2014-09-02 34