Self-Compression of Chain-of-Thought via Multi-Agent Reinforcement Learning
Proposes SCMA, a multi-agent RL framework, compresses reasoning chains by 39%, boosts accuracy by 10%.
Yiqun Chen, Jinyuan Feng, Wei Yang et al.
Proposes SCMA, a multi-agent RL framework, compresses reasoning chains by 39%, boosts accuracy by 10%.
Yiqun Chen, Jinyuan Feng, Wei Yang et al.
WebArbiter is a principle-guided reasoning reward model using structured text generation, surpassing scalar/template methods with 9.1% improvement on BoN accuracy.
Yao Zhang, Shijie Tang, Zeyu Li et al.
SCOUT framework decouples exploration and exploitation, enabling Qwen2.5-3B-Instruct to score 0.86 on unseen tasks, saving 60% GPU hours.
Haoyu Wang, Guozheng Ma, Shugang Cui et al.
Introduces Temporal Guidance (TeGu), leveraging temporal contrast to enhance LLM generation quality, achieving a 3.03% improvement on GSM8K.
Hong-Kai Zheng, Piji Li
CoFreeVLA reduces dual-arm self-collision rates from 0.54 to 0.23 and improves task success rates to 0.61 using a Vision-Language-Action model with risk estimation.
Yaohua Liu, Binkai Ou, Hengjun Zhang
UDBM introduces an uncertainty-guided diffusion bridge for single-step multi-task image restoration, outperforming state-of-the-art methods with 6× faster inference.
Luwei Tu, Jiawei Wu, Xing Luo et al.
CORDS maps discrete objects to continuous fields, enabling exact decoding of variable-sized sets with high accuracy.
Tin Hadži Veljković, Erik Bekkers, Michael Tiemann et al.
KAPSO employs a knowledge-grounded long-horizon optimization framework integrating git isolation, structured knowledge graphs, and episodic memory to enhance autonomous program synthesis.
Alireza Nadafian, Alireza Mohammadshahi, Majid Yazdani
MultiModal Fine-tuning with Synthetic Captions significantly improves performance on 13 benchmarks, especially in few-shot learning.
Shohei Enomoto, Shin'ya Yamaguchi
Dual-Stance Cooperative Debate (DSCD-Nav) enhances object navigation via multi-round evidence cross-checking, boosting success rate by 20% on HM3Dv2.
Weitao An, Qi Liu, Chenghao Xu et al.
MPF-Net detects high-fidelity AI video forgeries via hierarchical manifold deviation and micro-temporal fluctuation analysis, achieving 99.97% accuracy on VidProM.
Xinan He, Kaiqing Lin, Yue Zhou et al.
PTQ4ARVG framework achieves 6-bit quantization for ARVG models while maintaining competitive performance.
Xuewen Liu, Zhikai Li, Jing Zhang et al.
MAD method reduces cross-modal hallucinations by adaptive decoding, achieving 7.8% and 2.0% improvements on CMM and AVHBench.
Sangyun Chung, Se Yeon Kim, Youngchae Chee et al.
Proposes a vRKHS-based regularization framework for vector-valued regression under covariate shift, achieving optimal convergence rates.
Markus Holzleitner, Sergiy Pereverzyev, Sergei V. Pereverzyev et al.
Introduces Distributional Active Inference (DAIF), integrating active inference into distributional RL for model-free, efficient control in complex environments.
Abdullah Akgül, Gulcin Baykal, Manuel Haußmann et al.
Introduces ToxSearch-S with unsupervised speciation, boosting peak toxicity to 0.73 and semantic diversity, outperforming baseline in adversarial prompt search.
Onkar Shelar, Travis Desell
This paper introduces Persona Prompting (PP) to analyze its impact on social reasoning in LLMs, focusing on bias, rationale quality, and task performance using hate speech datasets.
Jing Yang, Moritz Hechtbauer, Elisabeth Khalilov et al.
Introduced a contextual runtime monitor framework, significantly enhancing safety in autonomous driving.
Alejandro Luque-Cerpa, Mengyuan Wang, Emil Carlsson et al.
DeepSeek-OCR 2 uses DeepEncoder V2 for visual causal flow, achieving a 3.73% performance boost on OmniDocBench v1.5.
Haoran Wei, Yaofeng Sun, Yukun Li
Proposed LoRA-based fine-tuning boosts 4-bit quantized LLM sensitivity awareness by 21.7%.
Dren Fazlija, Iyiola E. Olatunji, Daniel Kudenko et al.