Distilling Physical Priors into Streaming World Models
PhyS framework improves streaming world models' physical consistency by distilling physical priors, enhancing benchmarks like PhysicsIQ.
Liangliang Zhao, Junying Wang, Danni Yang et al.
PhyS framework improves streaming world models' physical consistency by distilling physical priors, enhancing benchmarks like PhysicsIQ.
Liangliang Zhao, Junying Wang, Danni Yang et al.
FlexSplat introduces a calibration-free 3D Gaussian splatting framework using joint depth and camera prediction, achieving near state-of-the-art results without known poses.
Amir Sabbaghziarani, Hanting Ye, Maria Gorlatova et al.
Proposes ICQ task, builds ISYV dataset and model, achieving 57% accuracy on person-centric video reasoning, improving long-horizon and cross-domain identification.
Shibo Gao, Chongxiao Wang, Chenglong Huang et al.
UniJEPA unifies image and video predictive modeling in a single latent space, using a single end-to-end loss with Gaussian regularization, enabling task-agnostic world modeling and zero-shot planning.
An Lanji, Dawei Liu, Jin Li et al.
GeoDistill-Refine employs multi-prompt fusion and signed distance fields, boosting annotation-free spacecraft segmentation IoU by 4.56% and Boundary F1 by 13.80%.
Yonglong Zhang, Zongwu Xie, Yang Liu
CANIS uses a generation-assisted semantic bridge for category-agnostic 3D canonicalization, improving classification and segmentation performance.
Kendong Liu, Yuxin Yao, Junhui Hou
DCMPC-Net uses cross-modal prompts for unified adverse-weather restoration, but the supplied text reports no numerical metrics.
Wanshu Fan, Yunzhe Zhang, Yue Shen et al.
HSMLA combines multi-scale linear attention with content-aware sparse softmax, achieving 4× speedup with high accuracy in high-res vision tasks.
Dong Liu, Yanxuan Yu, Renata Borovica-Gajic et al.
Gated Hindsight Distillation (GHD) leverages next screenshots as privileged training info, significantly improving mobile GUI agent success rates.
Weiwei Li, Junzhuo Liu, Tong Chu et al.
Vorch-Director matches residual corrections to flow-matching noise levels, improving long-horizon audio-visual stability; no numeric scores are provided.
Lisai Zhang, Yidi Wu, Qi Liu et al.
Proposes an iterative hybrid discrete-continuous viewpoint planning method to enhance UAV photogrammetry reconstruction accuracy and completeness.
Alan Grech, Daniel Pisani, Andre Grima et al.
Vorch-Streamer combines mixed Teacher Forcing and Diffusion Forcing with long-horizon Self Forcing, achieving 27.12 FPS real-time long-form audio-video generation.
Menglin Han, Yang Ding, Yulei Lu et al.
Vorch-IR is a unified multimodal video editing framework supporting multi-person identity and background replacement, built on LTX2 with attention mechanisms.
Yaole Wang, Xiaoyu Chen, Xin Ma et al.
Proposes CoCo-IR and TIE model, leveraging large multimodal models for multi-turn contextual image retrieval; achieves 39.4 mAP@5 (single-turn) and 44.1 R@1 (4-turn).
Shengcao Cao, Tanmaya Shekhar Dabral, Zhongli Ding et al.
AV-MSF combines multi-view images and few impact recordings to physically reconstruct object impact sounds, outperforming physics-based and data-driven baselines.
Zisen Shao, Zihao Wei, Derong Jin et al.
Leveraging foundation vision models for cross-species animal pose tracking, achieving high accuracy with limited labels.
Le Li, Daniela Ivanova, Nicolas Pugeault
MAGIC combines graph label propagation and geometric alignment to stabilize feature space in semi-supervised class-incremental learning, reducing drift and improving accuracy.
Yousef Abdi, Mohammad Asadpour, Yousef Seyfari
ToolArtist employs post-trained UMM with RL and RAD-GRPO to dynamically coordinate reasoning, tool use, and image generation, outperforming fixed pipeline methods.
Jiahao Zhao, Xiaomin Yu, Zhongxiang Sun et al.
Introduces CLIP-CC-Bench, a multi-model ensemble framework for evaluating paragraph-level video descriptions, covering 17 SOTA models with 200 movie clips.
Mukhtiar Ali, Harsh Dubey, Sugam Mishra et al.
DyLaR enhances video QA accuracy to 58.2% with under 20 tokens per query using dynamic latent reasoning.
Haotian Xia, Zilin Xiao, Junbo Zou et al.