Social Attention: Modeling Attention in Human Crowds
Social Attention learns nonlocal pedestrian importance, reducing mean ADE to 0.30 m and FDE to 2.59 m on ETH/UCY.
Anirudh Vemula, Katharina Muelling, Jean Oh
Social Attention learns nonlocal pedestrian importance, reducing mean ADE to 0.30 m and FDE to 2.59 m on ETH/UCY.
Anirudh Vemula, Katharina Muelling, Jean Oh
Unsupervised cross-lingual word mapping via adversarial training and Procrustes refinement surpasses supervised methods, achieving 66.2% accuracy on English-Italian translation.
Alexis Conneau, Guillaume Lample, Marc'Aurelio Ranzato et al.
Proposes polynomial approximation and Hermite polynomial-based estimators for asymptotically minimax estimation of L_r norms in Gaussian white noise models over Nikolskii-Besov spaces.
Yanjun Han, Jiantao Jiao, Rajarshi Mukherjee
Multi-agent self-play training induces emergent complex behaviors surpassing environment complexity.
Trapit Bansal, Jakub Pachocki, Szymon Sidor et al.
Proposes mixed precision training using FP16 storage, FP32 master weights, and loss scaling, achieving accuracy parity with FP32 while halving memory usage.
Paulius Micikevicius, Sharan Narang, Jonah Alben et al.
Proposes MLDG, a model-agnostic meta-learning method, achieving state-of-the-art domain generalization results on image classification and reinforcement tasks.
Da Li, Yongxin Yang, Yi-Zhe Song et al.
Proposes windowed chi-squared detector for sensor falsification, analyzing attack-induced system state deviations and comparing with static and CUSUM detectors.
Tunga R, Carlos Murguia, Justin Ruths
Proposes conditional imitation learning with high-level commands, enabling end-to-end autonomous driving with improved controllability and robustness.
Felipe Codevilla, Matthias Müller, Antonio López et al.
Gradual pruning method reduces large models by up to 10x with minimal accuracy loss, outperforming small dense models of same size.
Michael Zhu, Suyog Gupta
Proposed Deep Laplacian Pyramid Network for fast and accurate image super-resolution.
Wei-Sheng Lai, Jia-Bin Huang, Narendra Ahuja et al.
Construct minimally distorted adversarial examples via formal verification, enhancing adversarial training robustness by 4.2x.
Nicholas Carlini, Guy Katz, Clark Barrett et al.
HDLTex employs hierarchical deep neural networks to improve multi-level text classification, outperforming traditional methods with up to 97.97% accuracy.
Kamran Kowsari, Donald E. Brown, Mojtaba Heidarysafa et al.
Introduces FiLM: Feature-wise Linear Modulation, reducing error on CLEVR from 4.5% to 2.3%, surpassing prior state-of-the-art in visual reasoning.
Ethan Perez, Florian Strub, Harm de Vries et al.
AffordanceNet uses end-to-end deep learning for simultaneous object detection and pixel-level affordance recognition, achieving 150ms inference speed.
Thanh-Toan Do, Anh Nguyen, Ian Reid
Arabic MGB-3 Challenge uses i-vector features for dialect identification, achieving 75% accuracy.
Ahmed Ali, Stephan Vogel, Steve Renals
MuseGAN uses multi-track GANs to generate symbolic music, producing piano-rolls for five instruments.
Hao-Wen Dong, Wen-Yi Hsiao, Li-Chia Yang et al.
This paper systematically analyzes reproducibility issues in deep RL, highlighting the impact of randomness, environment variability, and implementation details.
Peter Henderson, Riashat Islam, Philip Bachman et al.
Proposed a novel estimator for mutual information in discrete-continuous mixtures, outperforming existing methods.
Weihao Gao, Sreeram Kannan, Sewoong Oh et al.
Matterport3D provides a large-scale indoor RGB-D panorama dataset with 10,800 views, enabling multi-task scene understanding including keypoint matching and semantic segmentation.
Angel Chang, Angela Dai, Thomas Funkhouser et al.
Unified graph embedding framework combining matrix factorization, random walks, and GNNs, achieving >85% accuracy in node classification.
William L. Hamilton, Rex Ying, Jure Leskovec