cs.CV 1506.09215

Unsupervised Learning from Narrated Instruction Videos

Proposes an unsupervised multimodal approach combining video and narration, achieving 85% accuracy in key task step detection on a new dataset.

Jean-Baptiste Alayrac, Piotr Bojanowski, Nishant Agrawal et al.

2015-07-01 70
cs.LG 1506.05439

Learning with a Wasserstein Loss

Proposes a Wasserstein distance-based loss for multi-label learning, utilizing entropic regularization for efficient approximation, enhancing semantic smoothness.

Charlie Frogner, Chiyuan Zhang, Hossein Mobahi et al.

2015-06-18 683 citations 51
math.NA 1506.03296

Randomized Iterative Methods for Linear Systems

Proposes a unified randomized iterative framework for linear systems, encompassing known algorithms like Kaczmarz and coordinate descent, with exponential convergence guarantees.

Robert M. Gower, Peter Richtárik

2015-06-10 33
stat.ML 1506.03134

Pointer Networks

Pointer Networks leverage attention as pointers, enabling variable-length input-output mapping, successfully applied to convex hull, Delaunay triangulation, and TSP.

Oriol Vinyals, Meire Fortunato, Navdeep Jaitly

2015-06-10 54
cs.CL 1505.06289

Text to 3D Scene Generation with Rich Lexical Grounding

Proposes a hybrid deep learning and rule-based approach for text-to-3D scene generation, achieving significant improvements in scene fidelity and diversity, with a new dataset and evaluation metrics.

Angel Chang, Will Monroe, Manolis Savva et al.

2015-05-23 123 citations 45
cs.CV 1504.08083

Fast R-CNN

Fast R-CNN achieves 9× faster training, 146× faster detection, with 66% mAP on VOC2007, by sharing features and end-to-end multi-task learning.

Ross Girshick

2015-04-30 54