PAVAS: Physics-Aware Video-to-Audio Synthesis
PAVAS integrates physical reasoning into video-to-audio synthesis, significantly enhancing physical consistency.
Oh Hyun-Bin, Yuhta Takida, Toshimitsu Uesaka et al.
PAVAS integrates physical reasoning into video-to-audio synthesis, significantly enhancing physical consistency.
Oh Hyun-Bin, Yuhta Takida, Toshimitsu Uesaka et al.
OneStory uses adaptive memory for multi-shot video, reaching 0.5813 T2MSV coherence.
Zhaochong An, Menglin Jia, Haonan Qiu et al.
Introduced NPC to optimize text-image alignment via negative prompts, achieving a 54% improvement on GenEval++.
Sangha Park, Eunji Kim, Yeongtak Oh et al.
SJD++ accelerates AR text-to-image generation by 2-3x using multi-token prediction and high-confidence reuse.
Yao Teng, Zhihuan Jiang, Han Shi et al.
RefBench-PRO improves localization accuracy in referring expression comprehension using Ref-R1 with Dynamic IoU-based GRPO.
Tianyi Gao, Hao Li, Han Fang et al.
CARD combines noise whitening and diffusion updates to address correlated noise, achieving state-of-the-art results on the CIN-D dataset.
Niki Nezakati, Arnab Ghosh, Amit Roy-Chowdhury et al.
DraCo integrates low-res sketches and multi-modal reasoning, boosting rare concept generation with +8% on GenEval and +0.91 on Imagine-Bench.
Dongzhi Jiang, Renrui Zhang, Haodong Li et al.
BulletTime employs 4D positional encoding and adaptive normalization to decouple scene dynamics from camera pose, enabling precise 4D control in video diffusion models.
Yiming Wang, Qihang Zhang, Shengqu Cai et al.
SFD asynchronously denoises semantics ahead of texture, reaching FID 1.04 on ImageNet 256×256 and accelerating convergence up to 100×.
Yueming Pan, Ruoyu Feng, Qi Dai et al.
LineAR compresses KV cache for efficient autoregressive image generation, reducing ImageNet FID from 2.77 to 2.68.
Ziran Qin, Youru Lv, Mingbao Lin et al.
COOPER integrates depth and segmentation modalities, achieving 6.91% improvement in spatial reasoning via a two-stage training framework.
Zefeng Zhang, Xiangzhao Hao, Hengzhu Tang et al.
Proposes KiMoI, a motion basis manipulation framework, to detect deepfake videos with subtle kinematic inconsistencies, improving generalization.
Alejandro Cobo, Roberto Valle, José Miguel Buenaposada et al.
PosterCopilot employs a three-stage training framework—PSFT, RL-VRA, RLAF—to improve layout accuracy and aesthetics, enabling layer-controlled iterative editing with high fidelity.
Jiazhe Wei, Ken Li, Tianyu Lao et al.
RELIC combines long-horizon memory with real-time streaming using compressed latent tokens, enabling interactive long-duration video world modeling at 16 FPS.
Yicong Hong, Yiqun Mei, Chongjian Ge et al.
Proposes test-time scaling with text embedding perturbation to boost diversity and quality in T2I diffusion models.
Hang Xu, Linjiang Huang, Feng Zhao
LAMP uses large language models for controllable video generation, enhancing motion control and user intent alignment.
Muhammed Burak Kizil, Enes Sanli, Niloy J. Mitra et al.
OneThinker unifies image and video reasoning across 10 tasks, trained on 600k data, with EMA-GRPO balancing rewards, achieving state-of-the-art results.
Kaituo Feng, Manyuan Zhang, Hongyu Li et al.
PPTArena benchmarks PowerPoint editing; PPTPilot, a structure-aware agent, surpasses VLMs by over 10% in complex, layout-sensitive tasks.
Michael Ofengenden, Yunze Man, Ziqi Pang et al.
Proposes MindGPT-4ov, a multi-stage post-training framework integrating data synthesis, curriculum fine-tuning, and reinforcement learning, achieving SOTA performance.
Wei Chen, Chaoqun Du, Feng Gu et al.
ReVSeg uses reinforcement learning to optimize explicit reasoning chains, achieving state-of-the-art video segmentation with interpretable decision trajectories.
Yifan Li, Yingda Yin, Lingting Zhu et al.