Moment Matching for Multi-Source Domain Adaptation
Proposed M3SDA method dynamically aligns feature distribution moments for multi-source domain adaptation, validated on DomainNet dataset.
Xingchao Peng, Qinxun Bai, Xide Xia et al.
Proposed M3SDA method dynamically aligns feature distribution moments for multi-source domain adaptation, validated on DomainNet dataset.
Xingchao Peng, Qinxun Bai, Xide Xia et al.
Generating high-fidelity images using Subscale Pixel Networks and Multidimensional Upscaling, achieving state-of-the-art results on CelebAHQ and ImageNet.
Jacob Menick, Nal Kalchbrenner
Introduced Touchdown dataset and model for natural language navigation and spatial reasoning in Google Street View, achieving 85.2% task success.
Howard Chen, Alane Suhr, Dipendra Misra et al.
3D human pose estimation in video using dilated temporal convolutions and semi-supervised training, reducing error by 11%.
Dario Pavllo, Christoph Feichtenhofer, David Grangier et al.
MeshNet introduces face-based features and mesh convolution for 3D shape representation, outperforming point cloud and voxel methods on ModelNet40.
Yutong Feng, Yifan Feng, Haoxuan You et al.
Implemented DDPG in TORCS for continuous control autonomous driving; achieved stable high-speed racing and overtaking in multiple scenarios.
Sen Wang, Daoyuan Jia, Xinshuo Weng
RCM combines reinforcement learning and self-supervised imitation to boost VLN, achieving 10% SPL improvement and better generalization.
Xin Wang, Qiuyuan Huang, Asli Celikyilmaz et al.
Proposed a self-paced adversarial training method, improving few-shot learning accuracy on CUB and Oxford-102 datasets.
Frederik Pahde, Oleksiy Ostapenko, Patrick Jähnichen et al.
Large-Scale Visual Active Learning with Deep Probabilistic Ensembles (DPE) reduces data needs, enhances performance.
Kashyap Chitta, Jose M. Alvarez, Adam Lesnikowski
Proposes heterogeneous auxiliary network feature mimicking, improving steering prediction MAE by 12.8%/52.1% on Udacity/Comma.ai datasets.
Yuenan Hou, Zheng Ma, Chunxiao Liu et al.
Introduces HDD, a multimodal driving dataset with novel annotation for driver behavior and causal reasoning, baseline models achieve over 80% mAP.
Vasili Ramanishka, Yi-Ting Chen, Teruhisa Misu et al.
Open Images V4 leverages large-scale multi-task annotations to advance image classification, detection, and relationship understanding.
Alina Kuznetsova, Hassan Rom, Neil Alldrin et al.
Proposes a face warping artifact detection method using CNNs, achieving 97.4% AUC on UADFV and 99.4% on DeepfakeTIMIT HQ, without relying on large fake datasets.
Yuezun Li, Siwei Lyu
Proposes Geometric Median-based filter pruning (FPGM), reducing over 52% FLOPs on ResNet-110 and 42% on ResNet-101 without accuracy loss.
Yang He, Ping Liu, Ziwei Wang et al.
Introduced RCN algorithm using relation networks for complex counting, excels on TallyQA dataset.
Manoj Acharya, Kushal Kafle, Christopher Kanan
Discrimination-aware channel pruning (DCP) integrates auxiliary losses to select channels with true discriminative power, achieving 30% channel reduction in ResNet-50 with a 0.39% accuracy boost on ImageNet.
Zhuangwei Zhuang, Mingkui Tan, Bohan Zhuang et al.
Structured Domain Randomization (SDR) integrates scene structure into synthetic data generation, significantly improving 2D vehicle detection on KITTI, outperforming traditional domain randomization.
Aayush Prakash, Shaad Boochoon, Mark Brophy et al.
SNIP: single-shot pruning based on connection sensitivity, pruning 80-99% connections at initialization with minimal accuracy loss.
Namhoon Lee, Thalaiyasingam Ajanthan, Philip H. S. Torr
Integrates multi-task perception (segmentation, depth) into end-to-end driving, boosting generalization by 20% in unseen environments and enabling accident explanation.
Zhihao Li, Toshiyuki Motoyoshi, Kazuma Sasaki et al.
RPNet estimates relative camera poses using deep learning, surpassing traditional methods by accurately recovering full translation vectors.
Sovann En, Alexis Lechervy, Frédéric Jurie