In-Context LoRA for Diffusion Transformers
IC-LoRA activates DiT in-context generation with 20–100 image sets, producing high-fidelity multi-image outputs without architectural changes.
Lianghua Huang, Wei Wang, Zhi-Fan Wu et al.
IC-LoRA activates DiT in-context generation with 20–100 image sets, producing high-fidelity multi-image outputs without architectural changes.
Lianghua Huang, Wei Wang, Zhi-Fan Wu et al.
Proposes MM-Det leveraging LMM-generated multi-modal forgery representations, achieving 92.0% AUC on DVF for diffusion video detection.
Xiufeng Song, Xiao Guo, Jiache Zhang et al.
Senna integrates LVLM with end-to-end models, reducing planning error by 27.12% and collision rate by 33.33%.
Bo Jiang, Shaoyu Chen, Bencheng Liao et al.
DynaMath evaluates the robustness of vision-language models in mathematical reasoning through dynamic question generation, revealing instability in handling variants.
Chengke Zou, Xingang Guo, Rui Yang et al.
LARP employs holistic queries and AR prior to achieve state-of-the-art FVD 57 in video generation.
Hanyu Wang, Saksham Suri, Yixuan Ren et al.
ProtoViT combines Vision Transformers for interpretable image classification, outperforming existing prototype models.
Chiyu Ma, Jon Donnelly, Wenjun Liu et al.
MoGe predicts 3D point clouds from single images using affine-invariant representations, enhanced by global and multi-scale local supervision, achieving state-of-the-art accuracy.
Ruicheng Wang, Sicheng Xu, Cassie Dai et al.
AVHBench is a benchmark for evaluating cross-modal hallucinations in audio-visual LLMs, revealing their limited understanding of complex relationships.
Kim Sung-Bin, Oh Hyun-Bin, JungMok Lee et al.
TranSPORTmer employs a unified transformer framework with Set Attention Blocks to perform trajectory forecasting, imputation, inference, and state classification, outperforming SOTA on sports datasets.
Guillem Capellera, Luis Ferraz, Antonio Rubio et al.
LongVU employs spatiotemporal adaptive compression using DINOv2 and cross-modal queries, enabling efficient long video understanding within limited context.
Xiaoqian Shen, Yunyang Xiong, Changsheng Zhao et al.
PyramidDrop reduces visual redundancy layer-wise, achieving 40% training time and 55% FLOPs inference acceleration with minimal performance loss.
Long Xing, Qidong Huang, Xiaoyi Dong et al.
Proposes Scene Language, a multimodal, program-based scene representation inferred from pre-trained models, enabling high-fidelity 3D/4D scene generation and editing without training.
Yunzhi Zhang, Zizhang Li, Matt Zhou et al.
TIPS combines spatial awareness via synthetic captions and self-supervised masked modeling, boosting dense image understanding.
Kevis-Kokitsi Maninis, Kaifeng Chen, Soham Ghosh et al.
DreamVideo-2 achieves zero-shot subject-driven video customization using reference attention and mask-guided motion, without test-time fine-tuning, outperforming SOTA on a new dataset.
Yujie Wei, Shiwei Zhang, Hangjie Yuan et al.
Introduces MC-Bench, a benchmark for evaluating multi-image visual grounding by 20+ models, highlighting performance gaps and future directions.
Yunqiu Xu, Linchao Zhu, Yi Yang
SANA employs deep compression autoencoder and linear DiT to generate 4096×4096 images efficiently on a single GPU, achieving 20× smaller size and 100× faster inference.
Enze Xie, Junsong Chen, Junyu Chen et al.
Mono-InternVL embeds visual experts into a pre-trained LLM with EViP, achieving superior multi-modal performance, surpassing 13 benchmarks.
Gen Luo, Xue Yang, Wenhan Dou et al.
SG-Nav uses online 3D scene graphs and hierarchical reasoning to boost zero-shot object navigation SR by over 10%.
Hang Yin, Xiuwei Xu, Zhenyu Wu et al.
Proposes MinorityPrompt, an online prompt optimization method enhancing low-likelihood sample generation in diffusion models.
Soobin Um, Jong Chul Ye
IterComp leverages multi-model preferences and iterative feedback to enhance compositional text-to-image generation, outperforming SOTA methods.
Xinchen Zhang, Ling Yang, Guohao Li et al.