cs.CV 2102.00719

Video Transformer Network

Proposed VTN, a transformer-based video recognition framework, trains 16x faster, runs 5x faster, with competitive accuracy on Kinetics-400.

Daniel Neimark, Omri Bar, Maya Zohar et al.

2021-02-01 31
cs.CV 2012.05901

Robust Consistent Video Depth Estimation

Proposed a new algorithm for estimating consistent depth maps and camera poses from monocular video, outperforming on the Sintel benchmark.

Johannes Kopf, Xuejian Rong, Jia-Bin Huang

2020-12-11 29
cs.CV 2011.14141

AdaBins: Depth Estimation using Adaptive Bins

AdaBins employs a transformer-based adaptive binning strategy, achieving state-of-the-art depth estimation with δ1=0.903 on NYU and RMS=0.364, surpassing previous methods.

Shariq Farooq Bhat, Ibraheem Alhashim, Peter Wonka

2020-11-28 1247 citations 38
cs.CV 2011.12948

Nerfies: Deformable Neural Radiance Fields

Nerfies extends NeRF with deformable volumetric fields, enabling photorealistic 3D reconstruction of dynamic scenes from casual videos, using coarse-to-fine optimization.

Keunhong Park, Utkarsh Sinha, Jonathan T. Barron et al.

2020-11-26 60