Face X-ray for More General Face Forgery Detection
Proposed Face X-ray detects face forgery by identifying blending boundaries, trained without fake images, achieving over 98% AUC on multiple datasets.
Lingzhi Li, Jianmin Bao, Ting Zhang et al.
Proposed Face X-ray detects face forgery by identifying blending boundaries, trained without fake images, achieving over 98% AUC on multiple datasets.
Lingzhi Li, Jianmin Bao, Ting Zhang et al.
Proposes TCoN with cross-domain co-attention for temporal feature alignment, achieving 15%+ accuracy boost in cross-domain action recognition.
Boxiao Pan, Zhangjie Cao, Ehsan Adeli et al.
Using Adversarial Representation Active Learning, achieve highest classification accuracy with minimal labels on datasets like MNIST.
Ali Mottaghi, Serena Yeung
Proposes P-CapsNets, removing routing, replacing convolution with capsule layers, and packaging tensors, achieving high efficiency with fewer parameters.
Zhenhua Chen, Xiwen Li, Chuhua Wang et al.
Proposes Differentiable Volumetric Rendering (DVR) for implicit 3D shape and texture learning from RGB images without 3D supervision.
Michael Niemeyer, Lars Mescheder, Michael Oechsle et al.
Proposed cascade cost volume improves high-res multi-view stereo accuracy by 35.6%, reducing GPU memory and runtime by over 50%.
Xiaodong Gu, Zhiwen Fan, Zuozhuo Dai et al.
Learn 6DoF closed-loop grasping from low-cost demonstrations using action-view rendering and Q-function for high success rates.
Shuran Song, Andy Zeng, Johnny Lee et al.
DeepDeform learns non-rigid RGB-D reconstruction from semi-supervised data: 390k frames, 5,533 pairs, and state-of-the-art accuracy.
Aljaž Božič, Michael Zollhöfer, Christian Theobalt et al.
Introduced Localized Narratives method, annotating 849k images by linking vision and language.
Jordi Pont-Tuset, Jasper Uijlings, Soravit Changpinyo et al.
Introduces CLOTH3D dataset and GCVAE model for realistic 3D clothed human generation, capturing garment topology and dynamics.
Hugo Bertiche, Meysam Madadi, Sergio Escalera
StarGAN v2 achieves diverse image synthesis across multiple domains, significantly outperforming baselines.
Yunjey Choi, Youngjung Uh, Jaejun Yoo et al.
ALFRED benchmark uses 25k natural language directives and expert demonstrations to challenge embodied vision-language models in complex household tasks.
Mohit Shridhar, Jesse Thomason, Daniel Gordon et al.
Redesigning normalization and removing progressive growing, the improved StyleGAN reduces artifacts, enhances image quality, and boosts invertibility, enabling higher resolution outputs.
Tero Karras, Samuli Laine, Miika Aittala et al.
PQ-NET generates 3D shapes via sequential part assembly, enabling autoencoding, interpolation, and single-view reconstruction.
Rundi Wu, Yixin Zhuang, Kai Xu et al.
Panoptic-DeepLab: A simple, strong, fast baseline for panoptic segmentation, achieving 65.5% PQ on Cityscapes.
Bowen Cheng, Maxwell D. Collins, Yukun Zhu et al.
BlendedMVS dataset combines 3D reconstruction and image blending, improving deep stereo model generalization.
Yao Yao, Zixin Luo, Shiwei Li et al.
Proposed an active learning method based on CNNs, significantly improving pedestrian detection accuracy.
Hamed H. Aghdam, Abel Gonzalez-Garcia, Joost van de Weijer et al.
Using unlabeled data in deep active learning improves image classification accuracy.
Oriane Siméoni, Mateusz Budnik, Yannis Avrithis et al.
MoCo introduces a large-scale contrastive learning framework with a momentum-updated encoder and a dynamic queue, achieving 60.6% top-1 accuracy on ImageNet with ResNet-50.
Kaiming He, Haoqi Fan, Yuxin Wu et al.
EVF employs fast experience encoding within a hierarchical Bayesian framework to enable rapid adaptation in visual prediction of novel objects, reducing prediction error by 15%.
Lin Yen-Chen, Maria Bauza, Phillip Isola