EfficientViT: Multi-Scale Linear Attention for High-Resolution Dense Prediction
EfficientViT achieves high-resolution dense prediction with multi-scale linear attention, reducing GPU latency by 13.9x.
Han Cai, Junyan Li, Muyan Hu et al.
EfficientViT achieves high-resolution dense prediction with multi-scale linear attention, reducing GPU latency by 13.9x.
Han Cai, Junyan Li, Muyan Hu et al.
Proposes 3DILG with irregular latent grids, boosting 3D shape reconstruction and generation performance.
Biao Zhang, Matthias Nießner, Peter Wonka
Imagen combines large-scale frozen T5-XXL encoder with diffusion models, achieving a state-of-the-art FID of 7.27 on COCO, greatly enhancing photorealism.
Chitwan Saharia, William Chan, Saurabh Saxena et al.
Proposes a deep transfer learning framework with a new taxonomy, analyzing source-target data relationships to improve image classification, especially in small-data scenarios.
Jo Plested, Musa Phiri, Tom Gedeon
Proposes label-invariant augmentation in representation space, generating hardest samples to improve semi-supervised graph classification.
Han Yue, Chunhui Zhang, Chuxu Zhang et al.
Region-aware metric learning with MCA achieves SOTA in open-world semantic segmentation, improving anomaly detection and few-shot incremental learning.
Hexin Dong, Zifan Chen, Mingze Yuan et al.
FRIH achieves fine-grained region-aware image harmonization with 38.19 dB PSNR on iHarmony4 dataset.
Jinlong Peng, Zekun Luo, Liang Liu et al.
SparseTT employs sparse Transformer attention to improve visual tracking accuracy, achieving 40FPS with 75% faster training than TransT.
Zhihong Fu, Zehua Fu, Qingjie Liu et al.
FixNet uses physical simulation and functional prediction to repair 3D objects, achieving 62.3% accuracy.
Yining Hong, Kaichun Mo, Li Yi et al.
Proposes Self-Calibrated Illumination (SCI) framework for fast, robust low-light image enhancement, outperforming SOTA with minimal parameters and high efficiency.
Long Ma, Tengyu Ma, Risheng Liu et al.
Proposes sim-2-sim transfer with VLN BERT, boosting VLN-CE success rate by 12%.
Jacob Krantz, Stefan Lee
ELEVATER is the first benchmark and toolkit for evaluating language-augmented visual models across 55 datasets with diverse efficiency metrics.
Chunyuan Li, Haotian Liu, Liunian Harold Li et al.
Self-Blended Images (SBI) enhances deepfake detection, achieving 99.64% AUC across datasets, improving cross-domain robustness.
Kaede Shiohara, Toshihiko Yamasaki
Proposes contrastive data collection to balance ArtEmis, improving emotional captioning with 20% CIDEr and 7% METEOR gains.
Youssef Mohamed, Faizan Farooq Khan, Kilichbek Haydarov et al.
Proposes a two-stage framework using CLIP latents and diffusion models to enhance image diversity and zero-shot control.
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol et al.
TopFormer employs a multi-scale Token Pyramid and scale-aware semantics, achieving 5% higher mIoU with low latency on mobile devices.
Wenqiang Zhang, Zilong Huang, Guozhong Luo et al.
NAFNet eliminates nonlinear activations, surpasses SOTA with 33.69dB PSNR on GoPro, using only 8.4% of the original computational cost.
Liangyu Chen, Xiaojie Chu, Xiangyu Zhang et al.
UniCL unifies contrastive learning in image-text-label space, achieving 14.5% improvement in zero-shot benchmarks.
Jianwei Yang, Chunyuan Li, Pengchuan Zhang et al.
Video diffusion model using 3D U-Net architecture enables high-quality long video synthesis with conditional sampling and joint training.
Jonathan Ho, Tim Salimans, Alexey Gritsenko et al.
Open-source benchmark framework for VG, analyzing impact of components on recall@N, model size, and efficiency.
Gabriele Berton, Riccardo Mereu, Gabriele Trivigno et al.