cs.CV 1511.02799

Neural Module Networks

Neural Module Networks combine deep learning with linguistic structure, achieving top results on VQA datasets.

Jacob Andreas, Marcus Rohrbach, Trevor Darrell et al.

2015-11-10 15
cs.CV 1509.01329

Semantic Amodal Segmentation

Proposes semantic amodal segmentation with large-scale datasets, achieving 78.5% AR on COCO, advancing scene understanding.

Yan Zhu, Yuandong Tian, Dimitris Mexatas et al.

2015-09-04 34
cs.CV 1506.09215

Unsupervised Learning from Narrated Instruction Videos

Proposes an unsupervised multimodal approach combining video and narration, achieving 85% accuracy in key task step detection on a new dataset.

Jean-Baptiste Alayrac, Piotr Bojanowski, Nishant Agrawal et al.

2015-07-01 76
cs.CV 1504.08083

Fast R-CNN

Fast R-CNN achieves 9× faster training, 146× faster detection, with 66% mAP on VOC2007, by sharing features and end-to-end multi-task learning.

Ross Girshick

2015-04-30 61