HELIOS: From midnight to noon, continuous outdoor urban scene relighting
HELIOS enables urban scene relighting from midnight to noon using unpaired datasets, enhancing structural consistency.
Hala Djeghim, Nathan Piasco, Luis Roldão et al.
HELIOS enables urban scene relighting from midnight to noon using unpaired datasets, enhancing structural consistency.
Hala Djeghim, Nathan Piasco, Luis Roldão et al.
Proposed a phase-aware attack framework achieving over 90% success rate on multimodal large language models.
Daizong Liu, Junhao Dong, Zhiyuan Ma et al.
StreamScout enhances timeline views progressively for efficient streaming video understanding, reducing token usage by 59%.
Ce Zhang, Jing Bi, Jinxi He et al.
BRF-GS employs 3D Gaussian splatting with hybrid BRDF kernels and spectral band selection for accurate multi-angle hyperspectral reflectance modeling.
Yiling Yao, Wenjuan Zhang, Bowen Wang et al.
DreamX-Creator employs Gated Cross-Modal Attention in a 7B model to achieve native 2K synchronized audio-video generation.
Jiashu Zhu, Yanhao Zheng, Ruitian Tian et al.
Proposes an audio-driven adversarial defense method to preserve visual fidelity in 3D talking face generation.
Rui-Qing Sun, Chen-Hao Cui, Hui-Yang Zhao et al.
LoRA fine-tuned OCR models convert handwritten pottery records to structured data, reducing error rates to below 1.5%.
Gissu Valentina Naghavi, Dominik Hagmann, Martin Kampel et al.
Proposes Discrete Diffusion Bridges (DDB) to resolve spatiotemporal misalignment in image generation using hybrid absorption and information-guided noise scheduling.
Xing Xie, Jiawei Liu, Shijun Zhou et al.
RIDGE selects frames in long videos using temporal signals, achieving top performance across four benchmarks.
Shanqing Xu, Meng Luo, Mengchen Qian et al.
3D-MRL enhances zero-shot 3D shape recognition accuracy to 50.9% using nested representation learning.
Márcus Lobo, Vitor Matias, Jeová Farias et al.
RAGDiffusion++ generates high-fidelity garment images using macro-retrieval and micro-alignment, reducing FID by 14.3%.
Yuhan Li, Xianfeng Tan, Fangao Zeng et al.
GoM framework enhances generalization in single-image multi-view synthesis using scene-disjoint validation and targeted diffusion adaptation, ranking first.
Jie Li, Xingchen Zou, Yuxuan Liang
Chat-Edit-3D++ enables flexible 3D and 4D scene editing via large language models, enhancing editing efficiency and outcomes.
Shuangkang Fang, Yufeng Wang, Yi-Hsuan Tsai et al.
First large-scale evaluation of text-guided models like Nano Banana for facial editing, assessing ~1M images.
Rahul Nair, Saurav Pandit, Hannah Kerner
This paper introduces GeBDA, a sequence prediction approach using the Gemma model for end-to-end building damage assessment from bi-temporal satellite images, achieving competitive localization and classification.
Olivier Dietrich, Krishna Sapkota, Konrad Schindler et al.
Utilizes pretrained video diffusion models for joint monocular depth and normal estimation via next-frame prediction, reducing data needs significantly.
Haosen Yang, Jifei Song, Zhensong Zhang et al.
State-conditioned memory router (LayerRecall) enhances long-range content recovery in video diffusion models, achieving top performance on MemoBench and MovieBench.
Yixuan Ding, Jiahao Kong, Wei Huang et al.
ARC-CT employs anatomy-guided contrastive learning, combining regional features to achieve 0.86 macro AUC on chest CT abnormalities.
Huseyin Umut Isik, Mehmet Alp Ozaydin, Sila Kurugol et al.
Scaling laws for video diffusion models in driving data show validation loss follows power laws; a 9B-parameter model predicts loss of ~0.0753 with 3.6% error.
Victor Besnier, Anh-Quan Cao, Elias Ramzi et al.
StreamEMS enhances video streaming understanding in vision-language models with a self-evolving memory scheme, excelling on OVO-Bench.
Yuxin Liu, Peiqin Zhuang, Yali Wang