Faster Training of Mask R-CNN by Focusing on Instance Boundaries
Introducing Edge Agreement Head accelerates Mask R-CNN training by 8.1%, improving mask boundary quality via edge gradient matching.
Roland S. Zimmermann, Julien N. Siems
Introducing Edge Agreement Head accelerates Mask R-CNN training by 8.1%, improving mask boundary quality via edge gradient matching.
Roland S. Zimmermann, Julien N. Siems
DANet achieves 81.5% mIoU on Cityscapes using dual attention mechanisms, enhancing scene segmentation accuracy.
Jun Fu, Jing Liu, Haijie Tian et al.
Joint autoregressive and hierarchical priors improve learned image compression, reducing file size by 15.8% over previous SOTA, outperforming BPG on PSNR and MS-SSIM.
David Minnen, Johannes Ballé, George Toderici
Proposed MLLC model uses latent context variables to improve video moment localization, validated on TEMPO dataset, outperforming prior methods.
Lisa Anne Hendricks, Oliver Wang, Eli Shechtman et al.
Soft Filter Pruning (SFP) allows pruned filters to be updated during training, maintaining model capacity and achieving over 42% FLOPs reduction with improved accuracy.
Yang He, Guoliang Kang, Xuanyi Dong et al.
Proposes a Holistic Scene Grammar (HSG) framework with MCMC optimization for single-image 3D scene parsing, improving layout and object detection accuracy by 20%.
Siyuan Huang, Siyuan Qi, Yixin Zhu et al.
A monocular video-based human avatar reconstruction method using SMPL model refinement, capturing facial, clothing, and fine details with 89.64% user preference.
Thiemo Alldieck, Marcus Magnor, Weipeng Xu et al.
Proposes a platform-aware neural architecture search (MNASNet) optimizing real-world latency and accuracy, achieving 75.2% Top-1 accuracy with 78ms latency on mobile.
Mingxing Tan, Bo Chen, Ruoming Pang et al.
Proposes a differentiable superpixel sampling network (SSN) combining deep features and soft clustering, outperforming traditional methods on benchmarks.
Varun Jampani, Deqing Sun, Ming-Yu Liu et al.
UNet++ employs nested dense skip connections, improving IoU by an average of 3.9 points over U-Net in medical image segmentation.
Zongwei Zhou, Md Mahfuzur Rahman Siddiquee, Nima Tajbakhsh et al.
CIRL combines imitation learning and deep reinforcement learning, achieving over 98% success in CARLA simulator for vision-based autonomous driving.
Xiaodan Liang, Tairui Wang, Luona Yang et al.
Proposes a self-supervised monocular endoscopy depth estimation method using SfM-generated sparse supervision, achieving submillimeter accuracy.
Xingtong Liu, Ayushi Sinha, Mathias Unberath et al.
Proposes VirtualHome, a simulation system that models household activities via programs, using deep learning to convert natural language and videos into executable task sequences, advancing robot task understanding.
Xavier Puig, Kevin Ra, Marko Boben et al.
Proposed a generative image inpainting system using gated convolution, achieving superior performance on Places2 and CelebA-HQ datasets.
Jiahui Yu, Zhe Lin, Jimei Yang et al.
FPN with ResNet50 for multi-class land segmentation achieved IoU 0.493 on DEEPGLOBE 2018.
Selim S. Seferbekov, Vladimir I. Iglovikov, Alexander V. Buslaev et al.
RedNet employs residual encoder-decoder with depth fusion and pyramid supervision, achieving 47.8% mIoU on SUN RGB-D.
Jindong Jiang, Lunan Zheng, Fei Luo et al.
AutoAugment uses reinforcement learning to automatically discover image augmentation policies, reducing error rates to 1.5% on CIFAR-10 and achieving 83.5% Top-1 accuracy on ImageNet.
Ekin D. Cubuk, Barret Zoph, Dandelion Mane et al.
AutoPruner integrates end-to-end training for filter pruning, achieving over 50% FLOPs reduction with less than 1.2% accuracy drop.
Jian-Hao Luo, Jianxin Wu
Detect image splicing via learned self-consistency using real photo datasets, achieving state-of-the-art performance.
Minyoung Huh, Andrew Liu, Andrew Owens et al.
Proposes NES-based black-box attack algorithms for query-limited, partial, and label-only models, achieving over 90% success on ImageNet and Google API with fewer queries.
Andrew Ilyas, Logan Engstrom, Anish Athalye et al.