Beyond Inpainting: Unleash 3D Understanding for Precise Camera-Controlled Video Generation
DepthDirector uses depth video guidance for precise camera control and consistent content generation.
Dong-Yu Chen, Yixin Guo, Shuojin Yang et al.
DepthDirector uses depth video guidance for precise camera control and consistent content generation.
Dong-Yu Chen, Yixin Guo, Shuojin Yang et al.
RAG-3DSG enhances 3D scene graphs by re-shot guided uncertainty estimation and retrieval-augmented generation, achieving state-of-the-art results in semantic consistency and precision.
Yue Chang, Rufeng Chen, Zhaofan Zhang et al.
Fast-ThinkAct compresses chain-of-thought reasoning into verbalizable latent representations, reducing inference latency by 89.3%.
Chi-Pin Huang, Yunze Man, Zhiding Yu et al.
Proposes the Semantic Lifecycle framework leveraging foundation models to enhance semantic acquisition, representation, and storage for embodied AI.
Shuai Chen, Hao Chen, Yuanchen Bei et al.
SpatialNav uses spatial scene graphs for zero-shot VLN, reaching 57.7% SR on R2R and 64.0% on R2R-CE.
Jiwen Zhang, Zejun Li, Siyuan Wang et al.
TAGRPO boosts GRPO in image-to-video generation via direct trajectory alignment, achieving significant reward improvements.
Jin Wang, Jianxiang Lu, Guangzheng Xu et al.
Mesh4D employs a VAE and diffusion model to reconstruct dynamic 3D meshes from monocular videos, capturing full shape and motion in a single pass.
Zeren Jiang, Chuanxia Zheng, Iro Laina et al.
DrivoR uses pretrained ViT and register tokens to compress multi-camera features, enabling efficient end-to-end autonomous driving with superior performance.
Ellington Kirby, Alexandre Boulch, Yihong Xu et al.
HuForDet employs dual-branch architecture combining RGB/frequency face analysis and multimodal semantic consistency, achieving SOTA detection with 90.22% AUC.
Xiao Guo, Jie Zhu, Anil Jain et al.
LocalDPO enhances video diffusion models by optimizing localized details, outperforming on Wan2.1 and CogVideoX.
Zitong Huang, Kaidong Zhang, Yukang Ding et al.
ResTok introduces hierarchical residual visual tokenizer, achieving gFID 2.34 with only 9 steps on ImageNet-256.
Xu Zhang, Cheng Da, Huan Yang et al.
Introduced SiT-Bench, a pure-text spatial reasoning benchmark with 3800+ questions, revealing a significant gap in global consistency for LLMs.
Zhongbin Guo, Zhen Yang, Yushan Li et al.
LTX-2 employs an asymmetric dual-stream Transformer with 14B (video) and 5B (audio) parameters, integrating cross-modal attention for synchronized audiovisual generation.
Yoav HaCohen, Benny Brazowski, Nisan Chiprut et al.
VLM4VLA uses minimal parameter fine-tuning to convert general VLMs into robot control policies, outperforming complex architectures.
Jianke Zhang, Xiaoyu Chen, Qiuyue Wang et al.
DGA-Net enhances camouflaged object detection using depth prompting and graph-anchor guidance, outperforming existing methods.
Yuetong Li, Qing Zhang, Yilin Zhao et al.
ReToK integrates redundant token padding and hierarchical semantic regularization, significantly enhancing autoregressive image generation for long sequences.
Zixuan Fu, Lanqing Guo, Chong Wang et al.
Fusion-SSAT leverages feature fusion of self-supervised reconstruction and global classification, achieving state-of-the-art cross-dataset deepfake detection with AUC 0.9613.
Shukesh Reddy, Srijan Das, Abhijit Das
SpaceTimePilot model achieves generative rendering of dynamic scenes with independent control over camera viewpoint and motion sequence.
Zhening Huang, Hyeonho Jeong, Xuelin Chen et al.
PhyGDPO enhances text-to-video generation's physical consistency via physics-guided preference optimization, outperforming existing methods.
Yuanhao Cai, Kunpeng Li, Menglin Jia et al.
HY-Motion 1.0 scales DiT-based flow matching to 1B+ parameters, enabling high-fidelity text-to-3D human motion generation.
Yuxin Wen, Qing Shuai, Di Kang et al.