Ref-NeRF: Structured View-Dependent Appearance for Neural Radiance Fields
Ref-NeRF reparameterizes view-dependent radiance using reflection directions, greatly improving specular reflection realism.
Dor Verbin, Peter Hedman, Ben Mildenhall et al.
Ref-NeRF reparameterizes view-dependent radiance using reflection directions, greatly improving specular reflection realism.
Dor Verbin, Peter Hedman, Ben Mildenhall et al.
Introduces KR distance and barycenter based on unbalanced optimal transport, enabling comparison of measures with different total mass.
Florian Heinemann, Marcel Klatt, Axel Munk
LiRA, a likelihood ratio attack based on first principles, improves low-FPR true positive rate by 10×, challenging existing privacy defenses.
Nicholas Carlini, Steve Chien, Milad Nasr et al.
ProtoPool model achieves interpretable image classification via differentiable prototype assignment, excelling on CUB-200-2011 and Stanford Cars datasets.
Dawid Rymarczyk, Łukasz Struski, Michał Górszczak et al.
Active inference using variational Bayesian inference enhances robot state estimation and control robustness under uncertainty.
Pablo Lanillos, Cristian Meo, Corrado Pezzato et al.
Mask2Former, a Transformer-based universal segmentation architecture, achieves SOTA on four datasets, outperforming specialized models with a novel masked attention mechanism.
Bowen Cheng, Ishan Misra, Alexander G. Schwing et al.
MViTv2 employs decomposed relative positional embeddings and residual pooling to enhance multiscale vision transformers, achieving 88.8% ImageNet accuracy and superior detection and recognition performance.
Yanghao Li, Chao-Yuan Wu, Haoqi Fan et al.
Dream Fields combines NeRF and CLIP for zero-shot text-guided 3D object generation, achieving 68.3% R-Precision without 3D supervision.
Ajay Jain, Ben Mildenhall, Jonathan T. Barron et al.
This paper compares prompting, imitation learning, and preference modeling across model scales, highlighting preference ranking's advantages and sample efficiency improvements via pretraining.
Amanda Askell, Yuntao Bai, Anna Chen et al.
MonoScene introduces monocular 3D semantic scene completion using a dual 2D/3D UNet framework with FLoSP and CRP modules, outperforming existing methods.
Anh-Quan Cao, Raoul de Charette
Diffusion Autoencoders combine meaningful high-level semantics with detailed reconstruction, enabling attribute manipulation and efficient representation learning.
Konpat Preechakul, Nattanat Chatthee, Suttisak Wizadwongsa et al.
VQ-Diffusion combines VQ-VAE and DDPM for text-to-image generation, achieving 15x speed improvement.
Shuyang Gu, Dong Chen, Jianmin Bao et al.
Blended Diffusion combines DDPM and CLIP for localized text-guided image editing, ensuring background preservation with high realism.
Omri Avrahami, Dani Lischinski, Ohad Fried
Introduces OOD-CV benchmark to evaluate vision models' robustness against individual nuisances like pose, texture, weather; shows current methods have limited improvements.
Bingchen Zhao, Shaozuo Yu, Wufei Ma et al.
GMFlow replaces local regression with global matching, reaching Sintel EPE 1.08 after one refinement, better than 31-step RAFT.
Haofei Xu, Jing Zhang, Jianfei Cai et al.
Proposes persistent homology dimension (PHD) as a topological measure to estimate neural network intrinsic dimension and predict generalization error.
Tolga Birdal, Aaron Lou, Leonidas Guibas et al.
VIOLET introduces an end-to-end video-language Transformer with Masked Visual-token Modeling, achieving SOTA on multiple tasks.
Tsu-Jui Fu, Linjie Li, Zhe Gan et al.
EvDistill combines bidirectional reconstruction and cross-modal distillation, reaching 58.02% mIoU on unlabeled DDD17 events.
Lin Wang, Yujeong Chae, Sung-Hoon Yoon et al.
Proposes a data-driven multi-agent simulation framework enabling zero-shot transfer of autonomous driving policies, validated on real vehicles with high success rates.
Tsun-Hsuan Wang, Alexander Amini, Wilko Schwarting et al.
mip-NeRF 360 leverages non-linear scene parameterization and online distillation, reducing error by 57% for unbounded scenes with high realism.
Jonathan T. Barron, Ben Mildenhall, Dor Verbin et al.