Pano3D: Unified 3D Reconstruction and Panoptic Segmentation
Pano3D achieves state-of-the-art 3D panoptic segmentation performance on ScanNet dataset.
Victor Barberteguy, Ahmet Iscen, Mathilde Caron et al.
Pano3D achieves state-of-the-art 3D panoptic segmentation performance on ScanNet dataset.
Victor Barberteguy, Ahmet Iscen, Mathilde Caron et al.
MeEvo combines natural evolution and metacognitive reflection through cyclic alternation, significantly improving search stability and solution quality on complex optimization tasks.
Zishang Qiu, Xinan Chen, Rong Qu et al.
SkillMutator enhances LLM agent skill security detection to 88.2% across 13 attack categories.
Youngduk Kim, Minkyoo Song, Seungwon Shin
WAM4D integrates geometric priors via spatial register tokens, enabling fast, spatially consistent 4D world modeling for robot manipulation.
Ying Li, Xiaobao Wei, Jiajun Cao et al.
AlignADV combines DPO and behavioral fingerprints, cutting training steps by up to 40.6%.
Yuewen Mei, Tong Nie, Jie Sun et al.
RT-VLA employs multi-level knowledge distillation to compress SimLingo's capabilities into a real-time, efficient model, reducing inference time by 44.8× while maintaining performance.
Xiangyu Huang, Zhenlin Hua, Han Zhou et al.
Proposes a co-evolutionary SNN ensemble framework based on marginal contribution fitness, significantly improving multi-task performance.
Catherine Rodriquez, James Ghawaly
InterleaveThinker employs a multi-agent framework with a planner and critic, achieving high-quality interleaved text-image generation with step-wise reinforcement learning, improving performance on benchmarks by over 50%.
Dian Zheng, Harry Lee, Manyuan Zhang et al.
Flow Reversal Steering (FRS) leverages reverse flow models to map coarse actions into high-quality behaviors, boosting zero-shot control and rapid learning in robotic policies.
Andy Tang, William Chen, Andrew Wagenmaker et al.
Proposes Modality Forcing, a post-training method enabling a single DiT model to jointly generate image and sparse depth data, achieving 57% reduction in AbsRel and scaling with model size.
Bardienus Pieter Duisterhof, Deva Ramanan, Jeffrey Ichnowski et al.
SpatialClaw employs code as an action interface, achieving 59.9% average accuracy across 20 spatial reasoning benchmarks, outperforming recent models by 11.2%.
Seokju Cho, Ryo Hachiuma, Abhishek Badki et al.
Using large language models (e.g., Claude 4.7) for automated reproducibility assessment in social sciences, matching effect sizes within ±0.05 and supporting conclusions with high accuracy.
Tobias Holtdirk, Pietro Marcolongo, Anna Steinberg Schulten et al.
This paper analyzes the sparsity and geometric structure of on-policy distillation (OPD), revealing small, coordinate-sparse updates that are spectrally concentrated and deviate from source principal directions.
Guo Yu, Wenlin Liu, Yulan Hu et al.
Flex4DHuman employs relative camera-pose encoding within a diffusion framework to synthesize synchronized multi-view videos from monocular or sparse inputs, surpassing prior methods without explicit geometry priors.
Jen-Hao Cheng, Yipeng Wang, Hao Zhang et al.
Introduces operads as a formal framework for question decomposition, with operadic consistency correlating strongly with model accuracy across multiple datasets.
Nathaniel Bottman, Kyle Richardson
Proposed MCR-Bionic hand integrates anatomical structural priors with hydraulic artificial muscles, achieving dexterous manipulation via wrist-finger coupling and intrinsic muscle pathways.
Haosen Yang, Guowu Wei
EWSegNet combines spatial and spectral features for efficient waste segmentation in cluttered backgrounds, achieving high accuracy with low computational cost.
Mamoona Javaid, Mubashir Noman, Abdul Hannan et al.
VISA uses offline VLM auditing to improve 3D semantic occupancy mIoU, significantly enhancing rare-class performance.
Ruiqi Xian, Yuehan Xian, Jing Liang et al.
CQC-RAG introduces cross-query consistency to enhance robustness in retrieval-augmented generation, outperforming baselines by +4.76 EM on TriviaQA and +9.12 EM on MuSiQue.
Yanjia Sun, Sifan Liu, Jie Shao
Introduced Smoothed-KL weighting, validated on CIFAR-10 and CelebA-64 with 0.45 FID improvement on average.
Lei Li