Not All Noises Are Created Equally:Diffusion Noise Selection and Optimization
Proposes noise selection and optimization based on inversion stability, improving diffusion model outputs by up to 72.5%.
Zipeng Qi, Lichen Bai, Haoyi Xiong et al.
Proposes noise selection and optimization based on inversion stability, improving diffusion model outputs by up to 72.5%.
Zipeng Qi, Lichen Bai, Haoyi Xiong et al.
VD3D introduces Plücker coordinate-based spatiotemporal camera embeddings for controlling large video transformers, achieving state-of-the-art accuracy.
Sherwin Bahmani, Ivan Skorokhodov, Aliaksandr Siarohin et al.
VISA integrates large multimodal language models with segmentation decoders to enable reasoning-based video object segmentation, outperforming state-of-the-art methods.
Cilin Yan, Haochen Wang, Shilin Yan et al.
FANet employs the AFE module combining SCM and FRM to enhance multi-scale features, achieving state-of-the-art semantic segmentation in cluttered scenes.
Muhammad Ali, Mamoona Javaid, Mubashir Noman et al.
LightenDiffusion combines Retinex theory with diffusion models for unsupervised low-light image enhancement, outperforming existing methods.
Hai Jiang, Ao Luo, Xiaohong Liu et al.
PaliGemma combines SigLIP-So400m and Gemma-2B, with 3B parameters, demonstrating broad transfer capabilities across tasks.
Lucas Beyer, Andreas Steiner, André Susano Pinto et al.
Comprehensive review of embodied AI leveraging MLMs and WMs, focusing on perception, interaction, and sim-to-real transfer, with experimental benchmarks.
Yang Liu, Weixing Chen, Yongjie Bai et al.
Introducing EVI-MAE, combining video and IMU data, achieving 92.78% accuracy in action recognition.
Mingfang Zhang, Yifei Huang, Ruicong Liu et al.
Tailor3D leverages dual-side images and LoRA-based triplane Transformer for fast, personalized 3D asset editing, achieving high consistency and efficiency.
Zhangyang Qi, Yunhan Yang, Mengchen Zhang et al.
Proposes TSGi, integrating scene graphs with CLIP-based multimodal learning, achieving 57.77% balanced accuracy in traffic accident classification, a near 5% improvement.
Aaron Lohner, Francesco Compagno, Jonathan Francis et al.
GenArtist employs a multimodal large language model as an agent for unified image generation and editing, surpassing DALL-E 3 with over 7% improvements on benchmarks.
Zhenyu Wang, Aoxue Li, Zhenguo Li et al.
Test-time contrastive concepts improve open-vocabulary semantic segmentation by automatically generating query-specific negative samples, boosting accuracy.
Monika Wysoczańska, Antonin Vobecky, Amaia Cardiel et al.
OneRestore framework uses cross-attention to restore complex image degradations, improving PSNR and SSIM.
Yu Guo, Yuan Gao, Yuxu Lu et al.
Proposes Dy-DCA, a content-aware single-model super-resolution framework with dynamic routing and compiler optimization, achieving 33FPS on mobile with 1.7× speedup.
Gen Li, Zhihao Shu, Jie Ji et al.
GlyphDraw2 combines diffusion models and LLMs for automatic complex poster generation, supporting multilingual text and font control.
Jian Ma, Yonglin Deng, Chen Chen et al.
Introduced MMLongBench-Doc, a benchmark with 135 long PDFs, evaluating 14 LVLMs; top F1 score is 42.7%.
Yubo Ma, Yuhang Zang, Liangyu Chen et al.
FoleyCrafter integrates pre-trained text-to-audio models with semantic adapters and temporal controllers for high-quality, synchronized video sound synthesis.
Yiming Zhang, Yicheng Gu, Yanhong Zeng et al.
CORE4D combines real MoCap and synthetic retargeting to create 11K 4D human-object-human interaction sequences for collaborative object rearrangement.
Yun Liu, Chengwen Zhang, Ruofan Xing et al.
Proposes Dynamic Gaussian Marbles for monocular video view synthesis, combining isotropic Gaussian modeling, hierarchical merging, and priors to improve scene reconstruction.
Colton Stearns, Adam Harley, Mikaela Uy et al.
StableNormal reduces diffusion inference variance to produce stable, sharp surface normals without ensembling, excelling in complex scenes.
Chongjie Ye, Lingteng Qiu, Xiaodong Gu et al.