cs.CV 1411.4555

Show and Tell: A Neural Image Caption Generator

Proposes a deep neural image captioning model combining CNN and LSTM, achieving BLEU-4 of 27.7, outperforming previous methods.

Oriol Vinyals, Alexander Toshev, Samy Bengio et al.

2014-11-18 79
cs.CV 1411.4038

Fully Convolutional Networks for Semantic Segmentation

Proposes Fully Convolutional Networks (FCN) for pixel-level semantic segmentation, achieving 62.2% mean IU, surpassing previous SOTA with end-to-end training.

Jonathan Long, Evan Shelhamer, Trevor Darrell

2014-11-15 55
cs.CV 1409.4842

Going Deeper with Convolutions

Proposed Inception architecture employs multi-scale convolutions and dimension reduction, achieving 28.6% Top-5 error on ImageNet with 1/12 parameters of AlexNet.

Christian Szegedy, Wei Liu, Yangqing Jia et al.

2014-09-17 45
cs.CV 1409.0575

ImageNet Large Scale Visual Recognition Challenge

Deep CNNs like AlexNet, VGG, ResNet trained on 14 million images achieved top-5 error rates below 7%, revolutionizing large-scale image recognition.

Olga Russakovsky, Jia Deng, Hao Su et al.

2014-09-02 34
cs.CV 1405.0312

Microsoft COCO: Common Objects in Context

MS COCO dataset with 91 categories, 2.5 million instances, enhances scene understanding and precise localization.

Tsung-Yi Lin, Michael Maire, Serge Belongie et al.

2014-05-02 46
cs.CV 1402.4963

Vesselness via Multiple Scale Orientation Scores

Introduces multi-scale invertible orientation scores for vessel enhancement, significantly improving crossing and bifurcation detection in retinal images.

Julius Hannink, Remco Duits, Erik Bekkers

2014-02-20 45