Event-based Vision: A Survey
Proposed a deep depth estimation method using event streams, achieving 85% accuracy improvement and microsecond-level real-time processing.
Guillermo Gallego, Tobi Delbruck, Garrick Orchard et al.
Proposed a deep depth estimation method using event streams, achieving 85% accuracy improvement and microsecond-level real-time processing.
Guillermo Gallego, Tobi Delbruck, Garrick Orchard et al.
Proposes a recurrent neural network for event-to-video reconstruction, improving image quality by over 20%, enabling application of standard vision algorithms.
Henri Rebecq, René Ranftl, Vladlen Koltun et al.
ContactDB analyzes and predicts grasp contact via thermal imaging, featuring 3750 3D meshes and 375K frames of RGB-D+thermal images.
Samarth Brahmbhatt, Cusuh Ham, Charles C. Kemp et al.
GA-Net introduces semi-global and local guided aggregation layers, achieving state-of-the-art accuracy with significantly reduced computational cost.
Feihu Zhang, Victor Prisacariu, Ruigang Yang et al.
A face de-occlusion method using 3D Morphable Model and GAN, enhancing face reconstruction accuracy.
Xiaowei Yuan, In Kyu Park
ObMan and a differentiable contact loss enable RGB hand-object reconstruction with improved physical grasp quality.
Yana Hasson, Gül Varol, Dimitrios Tzionas et al.
MVF-Net regresses 3DMMs from multiple views; its joint-loss model reaches MICC errors of 1.220 and 1.228.
Fanzi Wu, Linchao Bao, Yajing Chen et al.
This work introduces deep GCN architectures leveraging residual, dense, and dilated convolutions, training a 56-layer model that improves point cloud segmentation mIoU by 3.7%.
Guohao Li, Matthias Müller, Ali Thabet et al.
Fishyscapes benchmark evaluates anomaly detection in urban driving semantic segmentation, revealing blind spots in existing methods.
Hermann Blum, Paul-Edouard Sarlin, Juan Nieto et al.
VideoBERT models joint video and speech sequences using BERT with vector quantization, enabling high-level semantic understanding without supervision.
Chen Sun, Austin Myers, Carl Vondrick et al.
FCOS: anchor-free one-stage detector, AP 44.7%, simpler and faster than R-CNN variants.
Zhi Tian, Chunhua Shen, Hao Chen et al.
Habitat combines high-speed simulation and API for embodied AI, showing learned agents outperform SLAM in point-goal navigation with large experience scales.
Manolis Savva, Abhishek Kadian, Oleksandr Maksymets et al.
C2AE uses class-conditioned auto-encoders for open-set recognition, significantly outperforming existing methods on multiple datasets.
Poojan Oza, Vishal M Patel
HYPE measures human-perceived generative realism; truncated StyleGAN reached 27.6% HYPE∞ on FFHQ.
Sharon Zhou, Mitchell L. Gordon, Ranjay Krishna et al.
STM performs dense space-time memory reading for fast VOS, reaching J=88.7 on DAVIS-2016.
Seoung Wug Oh, Joon-Young Lee, Ning Xu et al.
Proposed Example Transfer Network (ETN) to mitigate negative transfer in Partial Domain Adaptation, achieving SOTA on benchmarks like Office-31.
Zhangjie Cao, Kaichao You, Mingsheng Long et al.
RNAN combines local and non-local attention; on Urban100 color denoising at σ=70, it reaches 27.45 dB.
Yulun Zhang, Kunpeng Li, Kai Li et al.
Introduces SPADE method to enhance semantic image synthesis with spatially-adaptive normalization for better visual fidelity and layout alignment.
Taesung Park, Ming-Yu Liu, Ting-Chun Wang et al.
Selective Kernel Networks (SKNet) enhance CNN performance via dynamic selection, outperforming existing models on ImageNet.
Xiang Li, Wenhai Wang, Xiaolin Hu et al.
Introduces a 3D pose generation model predicting human poses in indoor environments, outperforming existing methods.
Xueting Li, Sifei Liu, Kihwan Kim et al.