Face Forensics in the Wild
Constructed FFIW-10K dataset and proposed multi-instance attention model for multi-person face forgery detection.
Tianfei Zhou, Wenguan Wang, Zhiyuan Liang et al.
Constructed FFIW-10K dataset and proposed multi-instance attention model for multi-person face forgery detection.
Tianfei Zhou, Wenguan Wang, Zhiyuan Liang et al.
Extends NeRF to jointly encode geometry, appearance, and semantics, achieving high-quality 2D semantic labels with sparse scene annotations.
Shuaifeng Zhi, Tristan Laidlow, Stefan Leutenegger et al.
MVSNeRF combines multi-view geometry and neural rendering, reconstructing scenes from only three images with high speed and quality.
Anpei Chen, Zexiang Xu, Fuqiang Zhao et al.
Bakes NeRF into Sparse Neural Radiance Grid enabling 30FPS real-time rendering with less than 90MB storage.
Peter Hedman, Pratul P. Srinivasan, Ben Mildenhall et al.
Swin Transformer uses hierarchical design and shifted windows to achieve 87.3% top-1 accuracy on ImageNet and 58.7 AP on COCO detection.
Ze Liu, Yutong Lin, Yue Cao et al.
Proposes PlenOctrees for real-time NeRF rendering, achieving over 150 FPS, 3000x faster than traditional methods, with comparable quality.
Alex Yu, Ruilong Li, Matthew Tancik et al.
This survey reviews neural network quantization techniques, emphasizing low-bit (≤4 bits) methods for model compression and inference acceleration, covering algorithms, experiments, and future directions.
Amir Gholami, Sehoon Kim, Zhen Dong et al.
Proposes BCNet, a dual-layer GCN framework that significantly improves occlusion-aware instance segmentation, achieving +3.0 AP on COCO heavy occlusion scenarios.
Lei Ke, Yu-Wing Tai, Chi-Keung Tang
Using latent space regression to analyze GAN compositionality, enhancing image generation quality.
Lucy Chai, Jonas Wulff, Phillip Isola
Using Discrete Morse Theory (DMT) to train segmentation networks, significantly improving topological accuracy.
Xiaoling Hu, Yusu Wang, Li Fuxin et al.
GSVNet combines low-resolution flow and guided dynamic filtering, reaching 142 FPS and 71.8% mIoU on Cityscapes.
Shih-Po Lee, Si-Cun Chen, Wen-Hsiao Peng
Modular interactive VOS (MiVOS) decouples interaction-to-mask and propagation, using difference-aware fusion, achieving 85.2% IoU with fewer interactions.
Ho Kei Cheng, Yu-Wing Tai, Chi-Keung Tang
This study evaluates face obfuscation (blurring, overlay) on ImageNet recognition, showing less than 1% accuracy drop, supporting privacy-preserving vision.
Kaiyu Yang, Jacqueline Yau, Li Fei-Fei et al.
Repurposing GANs for one-shot semantic part segmentation achieves results comparable to supervised baselines.
Nontawat Tritrong, Pitchaporn Rewatbowornwong, Supasorn Suwajanakorn
Proposes ORE framework combining contrastive clustering and energy models for open world object detection, achieving state-of-the-art unknown detection and incremental learning.
K J Joseph, Salman Khan, Fahad Shahbaz Khan et al.
Proposed DyNeRF with space-time latent codes achieves high-quality 10s 30FPS multi-view video synthesis, reducing training time by 40x.
Tianye Li, Mira Slavcheva, Michael Zollhoefer et al.
Proposes a multi-attentional deepfake detection network reformulating the task as fine-grained classification, outperforming traditional binary classifiers.
Hanqing Zhao, Wenbo Zhou, Dongdong Chen et al.
P2-Net jointly describes and detects 2D-3D keypoints, achieving 97% matching recall on 7Scenes.
Bing Wang, Changhao Chen, Zhaopeng Cui et al.
IBRNet uses multi-view feature fusion and Transformer modules to synthesize high-res novel views without scene-specific optimization, achieving state-of-the-art generalization.
Qianqian Wang, Zhicheng Wang, Kyle Genova et al.
Transformers trained on 250M image-text pairs enable zero-shot text-to-image generation, outperforming previous models in diversity and realism.
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh et al.