A Unified Framework for Rethinking Policy Divergence Measures in GRPO
Unified clipping framework using KL3 estimator improves stability and exploration in GRPO, boosting math reasoning performance.
Qingyuan Wu, Yuhui Wang, Simon Sinong Zhan et al.
Unified clipping framework using KL3 estimator improves stability and exploration in GRPO, boosting math reasoning performance.
Qingyuan Wu, Yuhui Wang, Simon Sinong Zhan et al.
EmbedOpt optimizes protein diffusion model embeddings during inference, improving cryo-EM fitting and sparse distance constraints with fewer steps and enhanced robustness.
Minhuan Li, Jiequn Han, Pilar Cossio et al.
Proposes EntRGi for discrete diffusion language models, balancing gradient reliability and reward accuracy via entropy-based adaptive interpolation.
Atula Tejaswi, Litu Rout, Constantine Caramanis et al.
NeuroCanvas transforms multichannel EEG into images, improving seizure detection accuracy by 20% and reducing inference latency by 88%.
Yan Chen, Jie Peng, Moajjem Hossain Chowdhury et al.
RASA enhances MoE model safety by repairing safety-critical experts and maintaining routing consistency.
Jiacheng Liang, Yuhui Wang, Tanqiu Jiang et al.
EMA anchor and Top-k KL improve RL training for LLMs, achieving 53.9% accuracy on math reasoning.
Lunjun Zhang, Jimmy Ba
Supervised learning as lossy compression using finite blocklength analysis to reveal generalization and sample complexity.
Kosuke Sugiyama, Masato Uchida
AsymEP and Dyadic EP recover exact gradients in non-conservative systems; AsymEP reaches 94.9% on highly asymmetric MNIST networks.
Antonino Emanuele Scurria, Dimitri Vanden Abeele, Bortolo Matteo Mognetti et al.
Proposes basis rotation to mitigate gradient staleness in asynchronous pipeline parallelism, reducing training iterations by 81.7%.
Hyunji Jung, Sungbin Shin, Namhoon Lee
WGRPO pairs rare successes and failures, raising AIME 2025 Pass@8 from 16.8 to 22.2.
Yujuan Pang, Jiaxin Li, Xin Sheng et al.
LiDAR introduces a reward-guided sampling method using marginal samples and forward kernels, achieving 9.5× speedup over gradient guidance without neural backpropagation.
Yeongmin Kim, Donghyeok Shin, Byeonghu Na et al.
Introduces Contrastive Concept-Tree Search (CCTS), leveraging hierarchical semantic concepts to improve LLM-assisted algorithm discovery efficiency.
Timothee Leleu, Sudeera Gunathilaka, Federico Ghimenti et al.
RLAnything employs closed-loop optimization to dynamically shape environment, policy, and reward models, boosting performance on tasks like OSWorld by 9.1%.
Yinjie Wang, Tianbao Xie, Ke Shen et al.
CHASE framework leverages pretrained protein language models and flow matching to generate high-fitness protein variants efficiently without external predictors.
Amaru Caceres Arroyo, Lea Bogensperger, Ahmed Allam et al.
Proposes PPTP, a layer-based decoupling method, to enhance neural network privacy without sacrificing generalization.
Xingli Fang, Jung-Eun Kim
HopFormer uses head-specific n-hop sparse masking for structure injection, avoiding positional encodings, with linear complexity.
Sanggeon Yun, Raheeb Hassan, Ryozo Masukawa et al.
ECHO combines entropy and confidence-guided tree search with online pruning, significantly reducing collapse and bias in test-time RL, achieving +3.7% average gains on reasoning benchmarks.
Chu Zhao, Enneng Yang, Yuting Liu et al.
Proposes daVinci-Agency, leveraging GitHub PR chains for long-horizon data synthesis; fine-tuning on 239 samples yields 47% improvement on Toolathlon.
Mohan Jiang, Dayuan Fu, Junhao Shi et al.
Proposes Agentic Time Series Forecasting (ATSF) emphasizing perception, planning, action, reflection, and memory.
Mingyue Cheng, Xiaoyu Tao, Qi Liu et al.
PENCIL employs a plain Transformer with sampled local subgraphs for link prediction, achieving high accuracy with significantly fewer parameters.
Quang Truong, Yu Song, Donald Loveland et al.