HDMapNet: An Online HD Map Construction and Evaluation Framework
HDMapNet uses multi-modal sensor fusion for online high-definition map prediction, outperforming baselines by over 50%.
Qi Li, Yue Wang, Yilun Wang et al.
HDMapNet uses multi-modal sensor fusion for online high-definition map prediction, outperforming baselines by over 50%.
Qi Li, Yue Wang, Yilun Wang et al.
Proposed a post-training quantization algorithm for vision transformers, achieving 81.29% top-1 accuracy on ImageNet with DeiT-B model.
Zhenhua Liu, Yunhe Wang, Kai Han et al.
Video Swin Transformer uses local spatiotemporal attention, achieving 84.9% top-1 accuracy on Kinetics-400 with 28.2M parameters, outperforming global attention models.
Ze Liu, Jia Ning, Yue Cao et al.
This survey reviews 3D object detection methods for autonomous driving, emphasizing multi-modal fusion, datasets, and recent advances, with PV-RCNN achieving 82.86% mAP on KITTI.
Rui Qian, Xin Lai, Xirong Li
NeuS introduces a bias-free volume rendering approach combined with neural signed distance functions (SDF) for high-fidelity multi-view surface reconstruction, outperforming IDR and NeRF.
Peng Wang, Lingjie Liu, Yuan Liu et al.
HR-NAS integrates lightweight transformers and multi-scale search space to optimize high-resolution dense prediction networks.
Mingyu Ding, Xiaochen Lian, Linjie Yang et al.
Introduces sparse MoE into Vision Transformer, achieving 90.35% on ImageNet with 15B parameters, halving inference compute compared to dense models.
Carlos Riquelme, Joan Puigcerver, Basil Mustafa et al.
SimpleView method excels on ModelNet40 with half the parameters of PointNet++.
Ankit Goyal, Hei Law, Bowei Liu et al.
STCN uses L2 similarity for broader memory coverage, reaching 85.4 J&F and 20.2 FPS on DAVIS 2017.
Ho Kei Cheng, Yu-Wing Tai, Chi-Keung Tang
RobustNav benchmarks embodied navigation robustness under visual and dynamics corruptions, revealing significant performance drops and highlighting the need for improved adaptation.
Prithvijit Chattopadhyay, Judy Hoffman, Roozbeh Mottaghi et al.
Transferring 2D pretrained models to 3D point-cloud understanding via weight inflation and minimal fine-tuning achieves state-of-the-art results.
Chenfeng Xu, Shijia Yang, Tomer Galanti et al.
Patch Slimming removes redundant ViT patches, cutting DeiT-Ti FLOPs by 53.8% with only a 0.1-point ImageNet Top-1 drop.
Yehui Tang, Kai Han, Yunhe Wang et al.
Proposes Cross Pseudo Supervision (CPS) for semi-supervised semantic segmentation, achieving state-of-the-art results on Cityscapes and PASCAL VOC 2012.
Xiaokang Chen, Yuhui Yuan, Gang Zeng et al.
Proposes Exemplar-Based Open-Set Panoptic Segmentation Network (EOPSN), achieving 37.7% PQ on COCO with unknown class detection, advancing open-world scene understanding.
Jaedong Hwang, Seoung Wug Oh, Joon-Young Lee et al.
NExT-QA introduces a multi-task VideoQA benchmark focusing on causal and temporal reasoning, revealing models' weak performance in deep understanding.
Junbin Xiao, Xindi Shang, Angela Yao et al.
VPN++ employs dual-level knowledge distillation to fuse RGB videos and 3D poses, achieving fast, robust activity recognition without pose input at inference.
Srijan Das, Rui Dai, Di Yang et al.
DCT-NeRF learns trajectory fields for stable dynamic novel view synthesis.
Chaoyang Wang, Ben Eckart, Simon Lucey et al.
This study reveals that latent space similarities in ProtoPNet can be misleading under noise, with experiments showing JPEG compression and adversarial perturbations cause interpretability failures.
Adrian Hoffmann, Claudio Fanconi, Rahul Rade et al.
Proposes a deep learning-based real-time 3D human model with dynamic shape, motion, and textures learned from multi-view videos without detailed 3D supervision.
Marc Habermann, Lingjie Liu, Weipeng Xu et al.
GODIVA combines VQ-VAE and 3D sparse attention to generate open-domain videos; RM reaches 93.48 on MSR-VTT.
Chenfei Wu, Lun Huang, Qianxi Zhang et al.