On the Detection of Digital Face Manipulation
Attention-based CNN improves face forgery detection, achieving 99.76% AUC and better localization.
Hao Dang, Feng Liu, Joel Stehouwer et al.
Attention-based CNN improves face forgery detection, achieving 99.76% AUC and better localization.
Hao Dang, Feng Liu, Joel Stehouwer et al.
Proposes a dual-encoder neural network for simultaneous foreground and alpha matte estimation, outperforming state-of-the-art on synthetic and real datasets.
Qiqi Hou, Feng Liu
Proposes Dynamic Graph Attention (DGA) for multi-step visual reasoning, outperforming SOTA on benchmarks, enabling complex referring expression comprehension.
Sibei Yang, Guanbin Li, Yizhou Yu
FreiHAND dataset enhances 3D hand pose and shape estimation generalization using multi-view and semi-automated annotation.
Christian Zimmermann, Duygu Ceylan, Jimei Yang et al.
Proposes an end-to-end multi-modality MOT framework (mmMOT) integrating deep features of image and point cloud, boosting robustness and accuracy.
Wenwei Zhang, Hui Zhou, Shuyang Sun et al.
Proposes an end-to-end radar odometry system using learned feature embeddings, reducing errors by 68% and running ten times faster than state-of-the-art.
Dan Barnes, Rob Weston, Ingmar Posner
Holistic++ framework uses MCMC to optimize single-view 3D scene parsing and human pose estimation, significantly improving performance.
Yixin Chen, Siyuan Huang, Tao Yuan et al.
PlotQA with 28.9M QA pairs over 224k real-world plots advances complex reasoning in scientific visual question answering.
Nitesh Methani, Pritha Ganguly, Mitesh M. Khapra et al.
Introduces SPair-71k, a large-scale dataset with 70,958 image pairs across diverse viewpoints and scales, advancing semantic correspondence research.
Juhong Min, Jongmin Lee, Jean Ponce et al.
The PROX method uses 3D scene constraints to significantly reduce 3D human pose estimation errors.
Mohamed Hassan, Vasileios Choutas, Dimitrios Tzionas et al.
Proposed a hierarchical RL framework using subtask controllers and a meta controller to synthesize human-chair interactions, achieving a 31.61% success rate.
Yu-Wei Chao, Jimei Yang, Weifeng Chen et al.
AutoGAN optimizes GAN architecture via NAS, achieving an FID score of 12.42 on CIFAR-10.
Xinyu Gong, Shiyu Chang, Yifan Jiang et al.
Pixel2Mesh++ integrates multi-view perceptual features with graph convolutional deformation, achieving 80.30 F-score on ShapeNet, outperforming single-view methods.
Chao Wen, Yinda Zhang, Zhuwen Li et al.
Proposes MIND module with deep RL to enhance environment modeling and planning in EmbodiedQA, improving accuracy and generalization.
Juncheng Li, Siliang Tang, Fei Wu et al.
DIODE uses laser scanning to create high-density indoor-outdoor depth datasets, enabling cross-domain depth estimation.
Igor Vasiljevic, Nick Kolkin, Shanyi Zhang et al.
Structured3D dataset automatically extracts multi-structure geometric relations from professional interior designs, improving indoor scene understanding by 15% on benchmarks.
Jia Zheng, Junfei Zhang, Jing Li et al.
Output/feature distillation and encoder freezing raise incremental VOC segmentation mIoU to 71.6% without storing old images.
Umberto Michieli, Pietro Zanuttigh
Proposes InterFaceGAN, using linear hyperplanes in latent space for controllable face attribute editing.
Yujun Shen, Jinjin Gu, Xiaoou Tang et al.
Proposed multi-scale local planar guidance layers significantly enhance monocular depth estimation accuracy, outperforming existing methods on NYU Depth V2 and KITTI datasets.
Jin Han Lee, Myung-Kyu Han, Dong Wook Ko et al.
Proposes human-annotated frame shuffling to identify action classes requiring temporal cues, creating a 'Temporal Dataset' for benchmarking temporal modeling.
Laura Sevilla-Lara, Shengxin Zha, Zhicheng Yan et al.