VISA: VLM-Guided Instance Semantic Auditing for 3D Occupancy World Models
VISA uses offline VLM auditing to improve 3D semantic occupancy mIoU, significantly enhancing rare-class performance.
Ruiqi Xian, Yuehan Xian, Jing Liang et al.
VISA uses offline VLM auditing to improve 3D semantic occupancy mIoU, significantly enhancing rare-class performance.
Ruiqi Xian, Yuehan Xian, Jing Liang et al.
CQC-RAG introduces cross-query consistency to enhance robustness in retrieval-augmented generation, outperforming baselines by +4.76 EM on TriviaQA and +9.12 EM on MuSiQue.
Yanjia Sun, Sifan Liu, Jie Shao
Introduced Smoothed-KL weighting, validated on CIFAR-10 and CelebA-64 with 0.45 FID improvement on average.
Lei Li
SupraSNN employs a superscalar-inspired architecture with synapse-level parallelism, achieving 149μs latency and 0.025mJ/image on FPGA for MNIST, outperforming prior accelerators.
Seyed Sadra Ghavami, Mohammad Hossein Nikkhah, Mohammad Rasoul Roshanshah et al.
SkillCAT framework boosts LLM skill self-evolution by 49.69% through Contrastive Causal Extraction, Assessment-Augmented Evolution, and Topology-Aware Task Execution.
Kunfeng Chen, Qihuang Zhong, Juhua Liu et al.
Introduces Trajectory Sensitivity Score (TQS) based on dynamical systems stability to evaluate quantization impact on time-series models.
Mariya Pavlova, Harrison Bo Hua Zhu, Lidia Vitanova et al.
ProtoX-AD is a prototype-based self-explainable time series anomaly detection framework that achieves comparable performance to black-box models by leveraging transformation-aware latent representations.
Aitor Sánchez-Ferrera, Elisabeth Wetzer, Kristoffer Wickstrøm et al.
NTS-CoT reduces hallucinations in news timeline summarization using Chain-of-Thought, improving AR-1 by 23.4%.
Feng Lyu, Huiqin Yan, Sijing Duan et al.
FTP-1 is a generalist tactile policy improving contact-rich manipulation success by 31%.
Chengbo Yuan, Zicheng Zhang, Mingjie Zhou et al.
FCGRAFT uses function-level KV-cache grafting to improve embodied-policy success by 18.31% and synthesis speed by 2.3×.
Saehun Chun, Wonje Choi, Sera Choi et al.
Nous extracts behavioral parameters from trading data and attempts prompt-based injection to induce cognitive diversity, but results show limited effectiveness.
Haowei Qian
Introduces LocFac-RL using discrete diffusion models to enhance visual-textual reasoning efficiency, reducing computation by 26.9%.
Yoonjeon Kim, Yuhta Takida, Chieh-Hsin Lai et al.
DeepJEB++ leverages 2D latent space interpolation and foundation models to expand a small seed set into 15,360 labeled 3D jet engine brackets, with minimal resources.
Soyoung Yoo, Leekyo Jeong, Jinsu Ra et al.
This study compares Diff-in-Means (DiM) and INLP for extracting linear directions controlling model refusal, finding INLP's counterfactual flipping highly effective.
Elisabetta Rocchetti, Alfio Ferrara
This study systematically evaluates multimodal foundation models' perception, reasoning, and tool use in power defect detection, with detailed experimental data.
Quan Quan
RSMeM enhances remote sensing agents' tool usage accuracy by 6% through knowledge-enhanced memory evolution.
Bingxian Wu, Yu Zhang, Zonghao Guo et al.
Sparse2Act achieves cross-domain robot manipulation using action-aligned sparse 3D representations, with 86.9% success on LIBERO-10.
Yu Guo, Chang Yu, Siyu Ma et al.
VLADriveBench combines observational metrics and causal intervention to evaluate CoT–action causality in VLA autonomous driving models.
Thach Nguyen, Danhua Guo, Tom Lampo et al.
CAPED employs task-driven selective exposure, reducing incidental visual privacy leaks by over 70% while maintaining high task utility.
Siyu Shen, Fenghao Xu, Wenrui Diao et al.
This paper introduces DIRECT, a multimodal scene-aware routing framework that dynamically allocates test-time compute among embodied planners, reducing latency by up to 65% while maintaining or surpassing top performance.
Jadelynn Dao, Milan Ganai, Yasmina Abukhadra et al.