Pano3D: Unified 3D Reconstruction and Panoptic Segmentation
Pano3D achieves state-of-the-art 3D panoptic segmentation performance on ScanNet dataset.
Victor Barberteguy, Ahmet Iscen, Mathilde Caron et al.
Pano3D achieves state-of-the-art 3D panoptic segmentation performance on ScanNet dataset.
Victor Barberteguy, Ahmet Iscen, Mathilde Caron et al.
WAM4D integrates geometric priors via spatial register tokens, enabling fast, spatially consistent 4D world modeling for robot manipulation.
Ying Li, Xiaobao Wei, Jiajun Cao et al.
RT-VLA employs multi-level knowledge distillation to compress SimLingo's capabilities into a real-time, efficient model, reducing inference time by 44.8× while maintaining performance.
Xiangyu Huang, Zhenlin Hua, Han Zhou et al.
InterleaveThinker employs a multi-agent framework with a planner and critic, achieving high-quality interleaved text-image generation with step-wise reinforcement learning, improving performance on benchmarks by over 50%.
Dian Zheng, Harry Lee, Manyuan Zhang et al.
Proposes Modality Forcing, a post-training method enabling a single DiT model to jointly generate image and sparse depth data, achieving 57% reduction in AbsRel and scaling with model size.
Bardienus Pieter Duisterhof, Deva Ramanan, Jeffrey Ichnowski et al.
SpatialClaw employs code as an action interface, achieving 59.9% average accuracy across 20 spatial reasoning benchmarks, outperforming recent models by 11.2%.
Seokju Cho, Ryo Hachiuma, Abhishek Badki et al.
Flex4DHuman employs relative camera-pose encoding within a diffusion framework to synthesize synchronized multi-view videos from monocular or sparse inputs, surpassing prior methods without explicit geometry priors.
Jen-Hao Cheng, Yipeng Wang, Hao Zhang et al.
EWSegNet combines spatial and spectral features for efficient waste segmentation in cluttered backgrounds, achieving high accuracy with low computational cost.
Mamoona Javaid, Mubashir Noman, Abdul Hannan et al.
VISA uses offline VLM auditing to improve 3D semantic occupancy mIoU, significantly enhancing rare-class performance.
Ruiqi Xian, Yuehan Xian, Jing Liang et al.
Introduces LocFac-RL using discrete diffusion models to enhance visual-textual reasoning efficiency, reducing computation by 26.9%.
Yoonjeon Kim, Yuhta Takida, Chieh-Hsin Lai et al.
VLADriveBench combines observational metrics and causal intervention to evaluate CoT–action causality in VLA autonomous driving models.
Thach Nguyen, Danhua Guo, Tom Lampo et al.
VLGA introduces a dense 3D geometry expert supervised by LiDAR pointmap reconstruction, achieving state-of-the-art safety and driving scores in autonomous driving benchmarks.
Jin Yao, Dhruva Dixith Kurra, Tom Lampo et al.
Proposes Turbo-Inference, an iterative inference strategy leveraging detection-segmentation feedback, improving COCO and Cityscapes mAP by over 1% without retraining.
Zhen Zhao, Gang Zhang, Xiaolin Hu et al.
ExtremeWhenBench benchmark reveals that natural-language temporal grounding in long videos is a search problem; a retrieve-then-ground hybrid improves performance by 6.7x.
Sukmin Seo, Geewook Kim
Proposes ISAP-3D with explicit identity-slot alignment for stable 3D part generation, improving structural consistency.
Junlin Hao, Haoshuai Fu, Xibin Song et al.
This paper introduces a metadata-aware multi-prompt reasoning framework for zero-shot accident understanding, achieving a 15% improvement in harmonic mean score on CVPR benchmark.
Tarandeep Singh, Soumyanetra Pal, Soham Biswas et al.
BACON improves multimodal KV-cache compression by 7.5% on average, up to 30.9% under aggressive budgets.
Tianhao Chen, Yuheng Wu, Kelu Yao et al.
SG-PVR model enhances text-to-video generation semantic alignment using spatio-temporal scene graphs.
Hyomin Kim, Junghye Kim, Joanie Hayoun Chung et al.
UniReason-Med enhances 2D-to-3D medical VQA via a shared reasoning interface.
Mengzhuo Chen, Yan Shu, Chi Liu et al.
Next Forcing introduces multi-chunk prediction to accelerate training and improve accuracy in high-frame-rate video generation, achieving 94.1% success on RoboTwin.
Gangwei Xu, Qihang Zhang, Jiaming Zhou et al.