RbA: Segmenting Unknown Regions Rejected by All
Proposes RbA, a region-level outlier scoring method based on 'rejected by all,' improving unknown object segmentation with minimal supervision.
Nazir Nayal, Mısra Yavuz, João F. Henriques et al.
Proposes RbA, a region-level outlier scoring method based on 'rejected by all,' improving unknown object segmentation with minimal supervision.
Nazir Nayal, Mısra Yavuz, João F. Henriques et al.
CoMFormer leverages transformer architecture with adaptive distillation and pseudo-labeling to enable continual semantic and panoptic segmentation, outperforming existing methods with less forgetting.
Fabio Cermelli, Matthieu Cord, Arthur Douillard
Proposes Semantic Completion Learning (SCL) to enhance global-to-local cross-modal alignment, achieving SOTA results on vision-language benchmarks.
Yatai Ji, Rongcheng Tu, Jie Jiang et al.
Proposes LVDM, a latent diffusion model enabling high-fidelity long video synthesis, outperforming pixel-space models.
Yingqing He, Tianyu Yang, Yong Zhang et al.
Proposes a training-free 'Plug-and-Play' diffusion feature injection for text-guided image translation, preserving semantic layout.
Narek Tumanyan, Michal Geyer, Shai Bagon et al.
EDICT employs coupled affine transformations for exact diffusion inversion, reducing reconstruction error by over 50% compared to DDIM.
Bram Wallace, Akash Gokul, Nikhil Naik
InstructPix2Pix leverages GPT-3 and Stable Diffusion to generate 450K+ training pairs for instruction-based image editing.
Tim Brooks, Aleksander Holynski, Alexei A. Efros
Proposes Null-text inversion for high-fidelity real image editing using guided diffusion, achieving PSNR >30dB without model fine-tuning.
Ron Mokady, Amir Hertz, Kfir Aberman et al.
Latent-NeRF integrates shape and texture guidance via latent diffusion, enabling fast, controllable 3D generation.
Gal Metzer, Elad Richardson, Or Patashnik et al.
Unified dense correspondence model using Transformer surpasses SOTA in optical flow, stereo, and depth tasks with shared parameters and no cost volume.
Haofei Xu, Jing Zhang, Jianfei Cai et al.
Constructed 6.5TB DiffusionDB with 14M images and 1.8M prompts; analyzed prompt features, hyperparameters, and errors, revealing risks of misinformation and model bias.
Zijie J. Wang, Evan Montoya, David Munechika et al.
Proposes ID-unaware deepfake detection, reducing implicit identity leakage, improving cross-dataset generalization with Artifact Detection Module.
Shichao Dong, Jin Wang, Renhe Ji et al.
Introduces EMF metric and new dataset to evaluate monocular DVS, showing 1-2dB performance drop without multi-view cues.
Hang Gao, Ruilong Li, Shubham Tulsiani et al.
NeuPhysics learns 3D geometry and physics parameters from monocular videos, enabling editable dynamic scene reconstruction.
Yi-Ling Qiao, Alexander Gao, Ming C. Lin
Proposed HUMANISE dataset and scene-language conditioned generative model achieve diverse, semantically consistent 3D human motions, with key metrics showing 15% improvement over baselines.
Zan Wang, Yixin Chen, Tengyu Liu et al.
Imagic uses diffusion models for complex text-guided edits on a single real image.
Bahjat Kawar, Shiran Zada, Oran Lang et al.
Proposes weakly-supervised multi-granularity map learning, boosting indoor navigation success rate by 4%, through detailed environment representation and object localization.
Peihao Chen, Dongyu Ji, Kunyang Lin et al.
MAP introduces probabilistic distribution encoding for multimodal uncertainty, achieving SOTA on multiple vision-language tasks.
Yatai Ji, Junjie Wang, Yuan Gong et al.
Proposes mask-adapted CLIP with mask prompt tuning, achieving 29.6% mIoU on ADE20K-150, outperforming state-of-the-art by 8.5%.
Feng Liang, Bichen Wu, Xiaoliang Dai et al.
TEG-Track combines tactile and visual data to enhance 6D pose tracking of unseen objects, reducing rotation error by 30.9%.
Yun Liu, Xiaomeng Xu, Weihang Chen et al.