ActWorld: From Explorable to Interactive World Model via Action-Aware Memory
ActWorld adds action-aware memory to a 100K-video world model, enabling both navigation and mid-rollout object interaction.
Zhexiao Xiong, Yizhi Song, Hao Kang et al.
ActWorld adds action-aware memory to a 100K-video world model, enabling both navigation and mid-rollout object interaction.
Zhexiao Xiong, Yizhi Song, Hao Kang et al.
Pareto LoRA improves multimodal model balance by Pareto-optimal gradient integration, achieving 44.9% better image quality.
Xiwen Wei, Mark Nutter, Madhusudhanan Srinivasan et al.
Proposes Qwen-RobotWorld, a language-conditioned video world model using double-stream MMDiT and 8.6M embodied video-text pairs, achieving top performance on multiple benchmarks.
Jie Zhang, Xiaoyue Chen, Anzhe Chen et al.
MeshLoom is a feed-forward non-rigid mesh registration network that reconstructs vertex deformations across sequences within seconds, outperforming state-of-the-art methods.
Jianqi Chen, Jiraphon Yenphraphai, Xiangjun Tang et al.
Semantic Flip synthesizes out-of-distribution samples to improve embodied refusal, achieving an F1 of 0.9559 on spatial localization.
Dongbin Na, Chanwoo Kim, Giyun Choi et al.
GraphBEV++ integrates LocalAlign-v2 and GlobalAlign-v2 modules, systematically mitigating sensor calibration errors to enhance multi-modal perception robustness.
Ziying Song, Caiyan Jia, Lin Liu et al.
SelectStream enhances streaming video understanding with selective memory, achieving 82.67% on StreamingBench.
Haonan Ge, Yiwei Wang, Hang Wu et al.
GraphWorld leverages latent world models for end-to-end autonomous driving, reducing collision rates by 19.5% on nuScenes 6s horizon.
Ziying Song, Caiyan Jia, Lin Liu et al.
VinQA introduces two visual encoding methods, significantly improving visual citation accuracy in long-form multimodal document QA.
Young Rok Jang, Hyesoo Kong, Kyunghwan An et al.
Metis employs decoupled video generation and action prediction with Mixture-of-Transformers, achieving state-of-the-art autonomous driving performance.
Jingyu Li, Zhe Liu, Dongnan Hu et al.
Proposes TAG, converting egocentric videos into structured temporal graphs for zero-shot action recognition using VLMs, outperforming traditional pixel-based methods.
Bessie Dominguez-Dager, Francisco Gomez-Donoso, Miguel Cazorla et al.
Proposes a non-uniform timestep rescheduling method for diffusion inversion, reducing errors and improving image reconstruction accuracy by leveraging error analysis and dynamic programming.
Shangquan Sun, Ting Gong, Zhirui Liu et al.
CausalDrive employs a real-time causal autoregressive world model with flow-matching and self-distillation, achieving 12 FPS interactive autonomous driving simulation without future layout conditioning.
Tianyi Yan, Huan Zheng, Dubing Chen et al.
CoMET-Agent achieves conditional multi-event temporal grounding in long videos, improving [email protected] by 6.1%.
Yuanhao Zou, Arthad Kulkarni, Lucas Tonanez et al.
Spectral Forcing (SF) introduces a parameter-free, time-dependent 2D-DCT low-pass filter to improve pixel-space diffusion, boosting ImageNet FID by 14.5%.
Weichen Fan, Haiwen Diao, Penghao Wu et al.
MVEB benchmarks 23 tasks across 33 models, revealing diverse strengths and limitations in multi-modal video embeddings.
Adnan El Assadi, Roman Solomatin, Isaac Chung et al.
This study identifies a small set of attention heads—gaze heads—in VLMs that causally track the current description region, enabling effective inference-time control via attention masks.
Rohit Gandikota, David Bau
Introduces OmniVideo-100K, a large-scale dataset with structured scripts and evidence chains, boosting audio-visual reasoning by up to 20.59%.
Xinyue Cai, Chaoyou Fu, Yi-Fan Zhang et al.
Introducing RATS (Register Attention Transformers), which self-supervisedly discovers part-level structures with N learnable registers, achieving +12 mIoU on five segmentation benchmarks.
Timing Yang, Predrag Neskovic, Jansen Seheult et al.
Instruct-Particulate employs large-scale heterogeneous datasets and instruction-guided neural networks to efficiently predict 3D articulated structures, significantly improving generalization.
Ruining Li, Yuxin Yao, Matt Zhou et al.