UniGRPO: Unified Policy Optimization for Reasoning-Driven Visual Generation
UniGRPO optimizes text and image generation policies using GRPO, enhancing reasoning-driven visual generation quality.
Jie Liu, Zilyu Ye, Linxiao Yuan et al.
UniGRPO optimizes text and image generation policies using GRPO, enhancing reasoning-driven visual generation quality.
Jie Liu, Zilyu Ye, Linxiao Yuan et al.
DA-Flow combines diffusion and convolutional features to enhance optical flow estimation in degraded videos.
Jaewon Min, Jaeeun Lee, Yeji Choi et al.
WildWorld dataset offers over 450 actions and explicit state annotations for generative ARPG dynamic world modeling.
Zhen Li, Zian Meng, Shuwei Shi et al.
VISOR method enhances LVLM efficiency by sparsely selecting vision-language interactions, reducing inference cost.
Adrian Bulat, Alberto Baldrati, Ioannis Maniadis Metaxas et al.
AgentRVOS combines SAM3 and MLLM for zero-shot video object segmentation, achieving leading performance.
Woojeong Jin, Jaeho Lee, Heeseong Shin et al.
3DCity-LLM enhances 3D city-scale perception with a coarse-to-fine feature encoding strategy, leveraging a 1.2M-sample dataset.
Yiping Chen, Jinpeng Li, Wenyu Ke et al.
ABot-PhysWorld: 14B diffusion transformer for physics-aligned robotic manipulation video generation, trained on 3 million manipulation clips.
Yuzhi Chen, Ronghan Chen, Dongjie Huo et al.
TrajLoom combines Grid-Anchor Offset, VAE, and flow matching to predict 81-frame dense trajectories, improving stability and realism.
Zewei Zhang, Jia Jun Cheng Xian, Kaiwen Liu et al.
VideoDetective enhances long video understanding by integrating extrinsic query and intrinsic relevance, boosting VideoMME-long accuracy by 7.5%.
Ruoliu Yang, Chu Wu, Caifeng Shan et al.
UNITE achieves unified tokenization and latent diffusion with an autoencoder, reaching FID 2.12 on ImageNet.
Shivam Duggal, Xingjian Bai, Zongze Wu et al.
DualCoT-VLA enhances vision-language-action models with parallel reasoning for complex tasks, achieving state-of-the-art performance.
Zhide Zhong, Junfeng Li, Junjie He et al.
3D-Layout-R1 achieves language-guided spatial layout editing via scene graph reasoning, with a 15% IoU increase and 25% reduction in center-distance error.
Haoyu Zhen, Xiaolong Li, Yilin Zhao et al.
P-Flow employs test-time prompt optimization with VLMs to customize dynamic visual effects without model training, achieving high fidelity and diversity.
Rui Zhao, Mike Zheng Shou
GeoFusion-CAD extends long-sequence CAD generation via geometric state space, enhancing accuracy for 240-step commands.
Xiaolei Zhou, Chuangjie Fang, Jie Wu et al.
Proposes Adaptive Video Distillation with adaptive regression and temporal regularization, enabling stable few-step high-quality video synthesis.
Yuyang You, Yongzhi Li, Jiahui Li et al.
Proposes an image-conditioned RL framework for online VO parameter tuning, achieving 3x longer feature tracks and 3x lower computation.
Simone Nascivera, Leonard Bauersfeld, Jeff Delaune et al.
Proposes Predictive Regularization (PRe) to mitigate visual representation degradation in Multimodal Large Language Models, boosting visual fidelity and task performance.
Enguang Wang, Qiang Wang, Yuanchen Wu et al.
LumosX uses relational self-attention and cross-attention for personalized video generation, enhancing face-attribute alignment.
Jiazheng Xing, Fei Du, Hangjie Yuan et al.
VideoSeek actively seeks critical evidence using video logic flow, reducing frame usage by 93% and improving LVBench accuracy by 10.2 points.
Jingyang Lin, Jialian Wu, Jiang Liu et al.
LoD-Loc v3 employs instance silhouette alignment with synthetic data to improve urban aerial localization, achieving over 97% success in 2m/2° accuracy.
Shuaibang Peng, Juelin Zhu, Xia Li et al.