Video Transformer Network
Proposed VTN, a transformer-based video recognition framework, trains 16x faster, runs 5x faster, with competitive accuracy on Kinetics-400.
Daniel Neimark, Omri Bar, Maya Zohar et al.
Proposed VTN, a transformer-based video recognition framework, trains 16x faster, runs 5x faster, with competitive accuracy on Kinetics-400.
Daniel Neimark, Omri Bar, Maya Zohar et al.
LightLayers uses matrix decomposition to reduce parameters, achieving CIFAR-10 accuracy with only 1/3 of the original parameters.
Debesh Jha, Anis Yazidi, Michael A. Riegler et al.
Proposes a structured-light-based deep learning framework for endoscopic depth estimation, achieving an average error below 3mm on porcine datasets.
Max Allan, Jonathan Mcleod, Congcong Wang et al.
SETR uses a pure Transformer for segmentation, reaching 50.28% mIoU on ADE20K.
Sixiao Zheng, Jiachen Lu, Hengshuang Zhao et al.
STaR employs self-supervised neural radiance fields to reconstruct and track rigid objects in motion from multi-view videos, enabling photorealistic novel view synthesis without labels.
Wentao Yuan, Zhaoyang Lv, Tanner Schmidt et al.
Proposes a user-click guided, uncertainty-aware deep image matting framework that outperforms existing trimap-free methods and rivals trimap-based approaches.
Tianyi Wei, Dongdong Chen, Wenbo Zhou et al.
Proposed a new algorithm for estimating consistent depth maps and camera poses from monocular video, outperforming on the Sintel benchmark.
Johannes Kopf, Xuejian Rong, Jia-Bin Huang
Improved out-of-distribution detection in semantic segmentation using entropy maximization and meta classification, reducing error rate by 52%.
Robin Chan, Matthias Rottmann, Hanno Gottschalk
Neural Prototype Trees combine prototype learning and decision trees for interpretable fine-grained image recognition, excelling on the CUB-200-2011 dataset.
Meike Nauta, Ron van Bree, Christin Seifert
DiffusionNet employs a learnable diffusion layer for surface data propagation, achieving resolution and sampling invariance.
Nicholas Sharp, Souhaib Attaiki, Keenan Crane et al.
MaX-DeepLab uses mask transformers for end-to-end panoptic segmentation, achieving a 7.1% PQ increase on COCO.
Huiyu Wang, Yukun Zhu, Hartwig Adam et al.
Panoptic FCN achieves efficient panoptic segmentation with kernel generation, reaching 44.3% PQ on COCO.
Yanwei Li, Hengshuang Zhao, Xiaojuan Qi et al.
HybrIK combines analytical and neural inverse kinematics to improve 3D human pose and shape estimation, reducing MPJPE by 13.2mm on 3DPW.
Jiefeng Li, Chao Xu, Zhicun Chen et al.
AdaBins employs a transformer-based adaptive binning strategy, achieving state-of-the-art depth estimation with δ1=0.903 on NYU and RMS=0.364, surpassing previous methods.
Shariq Farooq Bhat, Ibraheem Alhashim, Peter Wonka
Proposes Neural Scene Flow Fields for monocular dynamic scene space-time view synthesis, outperforming existing methods with detailed 3D motion modeling.
Zhengqi Li, Simon Niklaus, Noah Snavely et al.
Nerfies extends NeRF with deformable volumetric fields, enabling photorealistic 3D reconstruction of dynamic scenes from casual videos, using coarse-to-fine optimization.
Keunhong Park, Utkarsh Sinha, Jonathan T. Barron et al.
MODNet is a real-time, trimap-free portrait matting model using objective decomposition, achieving 67fps.
Zhanghan Ke, Jiayu Sun, Kaican Li et al.
Semantic scene completion using local deep implicit functions, surpassing KITTI benchmark in geometric IoU.
Christoph B. Rist, David Emmerichs, Markus Enzweiler et al.
Autoencoder-based outlier detection with DBSCAN removes over 67% mislabeled images, improving dataset quality.
Yunhao Yang, Andrew Whinston
Hypersim leverages artist-created scenes to generate 77,400 photorealistic indoor images with detailed pixel labels, significantly advancing scene understanding.
Mike Roberts, Jason Ramapuram, Anurag Ranjan et al.