Point Transformer
Point Transformer employs multi-head attention and SortNet to achieve permutation-invariant point cloud features, capturing local and global structures.
Nico Engel, Vasileios Belagiannis, Klaus Dietmayer
Point Transformer employs multi-head attention and SortNet to achieve permutation-invariant point cloud features, capturing local and global structures.
Nico Engel, Vasileios Belagiannis, Klaus Dietmayer
Fusion of RGB and LiDAR data using 2D detectors and pixel mapping improves 3D object detection on nuScenes, achieving ~67% mAP.
Yilin Wang, Jiayi Ye
Proposes Vision Transformer (ViT), applying pure self-attention to image patches, achieving state-of-the-art results on ImageNet after large-scale pretraining.
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov et al.
SCOP introduces knockoff features as scientific control, achieving 57.8% parameter reduction and 60.2% FLOPs reduction on ResNet-101 with only 0.01% accuracy loss.
Yehui Tang, Yunhe Wang, Yixing Xu et al.
Proposes a self-supervised geometric feature discovery framework with interpretable attention, achieving SOTA vehicle re-identification performance.
Ming Li, Xinming Huang, Ziming Zhang
RxR introduces a multilingual VLN dataset with dense spatiotemporal grounding, significantly improving navigation accuracy and generalization.
Alexander Ku, Peter Anderson, Roma Patel et al.
NeRF++ introduces inverted sphere parameterization and dual scene modeling, significantly improving large-scale unbounded scene view synthesis.
Kai Zhang, Gernot Riegler, Noah Snavely et al.
Proposes AdAGeo combining attention and few-shot unsupervised domain adaptation, achieving over 13% improvement in cross-domain visual place recognition with only 5 target images.
Gabriele Moreno Berton, Valerio Paolicelli, Carlo Masone et al.
GRF learns a general radiance field from multi-view images, enabling high-quality 3D scene reconstruction and novel view synthesis.
Alex Trevithick, Bo Yang
This paper introduces RAPS, a conformal prediction method with regularization, significantly reducing prediction set size while guaranteeing coverage.
Anastasios Angelopoulos, Stephen Bates, Jitendra Malik et al.
PP-OCR is an ultra-lightweight OCR system with a 3.5M model for 6622 Chinese characters.
Yuning Du, Chenxia Li, Ruoyu Guo et al.
Introduces 3D-FUTURE, a large-scale dataset with 20,240 synthetic indoor images and 9,992 detailed textured furniture models for multi-task 3D scene understanding.
Huan Fu, Rongfei Jia, Lin Gao et al.
Network dissection systematically identifies semantic roles of individual units in CNNs and GANs, revealing object detectors and controllable scene features.
David Bau, Jun-Yan Zhu, Hendrik Strobelt et al.
AB3DMOT employs 3D Kalman filtering and Hungarian matching, achieving 207.4 FPS with state-of-the-art accuracy on KITTI.
Xinshuo Weng, Jianren Wang, David Held et al.
Lift-Splat-Shoot framework enables end-to-end bird’s-eye-view encoding from multi-view images via implicit 3D unprojection, improving perception and motion planning.
Jonah Philion, Sanja Fidler
Proposes a scene-agnostic neural view synthesis method using SfM and MVS, outperforming state-of-the-art on Tanks and Temples with over 50% LPIPS reduction.
Gernot Riegler, Vladlen Koltun
TransNet V2 combines multi-scale 3D features and frame similarities, achieving F1 scores of 77.9%, 96.2%, and 93.9%.
Tomáš Souček, Jakub Lokoč
I2L-MeshNet predicts 1D lixel heatmaps and reaches 55.7 mm MPJPE on Human3.6M for monocular 3D mesh recovery.
Gyeongsik Moon, Kyoung Mu Lee
NeRF-W introduces appearance and transient scene modeling, achieving 29.08 PSNR on landmark datasets, significantly improving 3D reconstructions from wild photo collections.
Ricardo Martin-Brualla, Noha Radwan, Mehdi S. M. Sajjadi et al.
pSp framework uses StyleGAN for image-to-image translation, directly embedding into W+ space.
Elad Richardson, Yuval Alaluf, Or Patashnik et al.