Trajectory Attention for Fine-grained Video Motion Control
Trajectory attention explicitly models pixel trajectories, boosting fine-grained camera motion control with improved long-range consistency.
Zeqi Xiao, Wenqi Ouyang, Yifan Zhou et al.
Trajectory attention explicitly models pixel trajectories, boosting fine-grained camera motion control with improved long-range consistency.
Zeqi Xiao, Wenqi Ouyang, Yifan Zhou et al.
FonTS employs a two-stage DiT pipeline with parameter-efficient fine-tuning and style adapters to achieve precise word-level typography and style control, significantly improving text rendering quality.
Wenda Shi, Yiren Song, Dengming Zhang et al.
AC3D leverages spectral analysis and model optimization to enable precise 3D camera control in video diffusion transformers, improving visual quality by 10% and training speed by 15%.
Sherwin Bahmani, Ivan Skorokhodov, Guocheng Qian et al.
Introduces VDMini, accelerating video diffusion models by pruning and consistency loss, achieving 2.5x speedup.
Yiming Wu, Zhenghao Chen, Huan Wang et al.
ChatRex enhances multimodal LLM perception with decoupled design, achieving 72.8% recall on COCO dataset.
Qing Jiang, Gen Luo, Yuqin Yang et al.
Type-R employs post-processing to automatically correct typos in text-to-image outputs, significantly improving text accuracy without degrading image quality.
Wataru Shimoda, Naoto Inoue, Daichi Haraguchi et al.
CHOICE benchmark systematically evaluates 23 remote sensing tasks for large vision-language models, revealing strengths and gaps in perception and reasoning, with 10,507 problems across 50 cities.
Xiao An, Jiaxing Sun, Zihan Gui et al.
This study conducts a hierarchical attention analysis in VLMs, revealing how middle layers facilitate cross-modal information flow and global feature storage.
Omri Kaduri, Shai Bagon, Tali Dekel
ConsisID introduces a tuning-free, frequency-based control scheme for identity-preserving text-to-video generation using DiT, achieving high fidelity without fine-tuning.
Shenghai Yuan, Jinfa Huang, Xianyi He et al.
PreF3R is a pose-free, feed-forward 3D Gaussian reconstruction method achieving 20FPS for real-time novel-view synthesis.
Zequn Chen, Jiezhi Yang, Heng Yang
This paper proposes a precise scaling law for video diffusion Transformers, integrating optimal hyperparameter prediction to enhance performance and reduce inference costs by 40.1%.
Yuanyang Yin, Yaqi Zhao, Mingwu Zheng et al.
EPS uses DCT-based spatial-temporal features for efficient patch sampling, reducing training data by up to 91.69% with 82.1× speedup.
Yiying Wei, Hadi Amirpour, Jong Hwan Ko et al.
Zoom Eye enhances multimodal LLMs' visual reasoning with tree search, boosting InternVL2.5-8B by 15.71% on HR-Bench.
Haozhan Shen, Kangjia Zhao, Tiancheng Zhao et al.
Introduces orthogonal subspace decomposition via SVD to enhance generalization in AI-generated image detection, outperforming SOTA with >94% AUC.
Zhiyuan Yan, Jiangming Wang, Peng Jin et al.
Using k-SAE to interpret diffusion model features reveals hierarchical semantic information, with intermediate layers capturing fine details.
Dahye Kim, Xavier Thomas, Deepti Ghadiyaram
Proposes truncated diffusion with multi-mode anchors, reducing steps from 20 to 2, achieving 88.1 PDMS at 45FPS for autonomous driving.
Bencheng Liao, Shaoyu Chen, Haoran Yin et al.
OminiControl leverages parameter reuse of DiT's VAE and Transformer, adding only 0.1% parameters, to achieve versatile multi-task image control surpassing specialized methods.
Zhenxiong Tan, Songhua Liu, Xingyi Yang et al.
Stable Flow finds vital layers in DiT to enable training-free, stable image editing.
Omri Avrahami, Or Patashnik, Ohad Fried et al.
DINO-X achieves state-of-the-art open-world object detection with 56.0 AP on COCO.
Tianhe Ren, Yihao Chen, Qing Jiang et al.
EAST combines dynamic ReLU, weight sharing, and cyclic sparsity to achieve 99.99% sparsity with competitive accuracy.
Andy Li, Aiden Durrant, Milan Markovic et al.