Distributional Soft Bellman Operator under the Cramér Geometry
Introduces Cramér geometry-based distributional soft Bellman operator, proving contraction and unique fixed point for RL policy evaluation.
Keru Wang, Yixin Deng, Yao Lyu et al.
Introduces Cramér geometry-based distributional soft Bellman operator, proving contraction and unique fixed point for RL policy evaluation.
Keru Wang, Yixin Deng, Yao Lyu et al.
The study proposes a one-step lowest-variance selection method using a Gaussian random-field model to analyze total correlation and square root collision threshold.
Linjun Li
MultiLoReFT decouples shared and modality-specific information via low-rank representation fine-tuning, improving prediction and interpretability.
Sana Tonekaboni, Viktoria Schuster, Caroline Uhler
Proposed a knowledge-dependent linear truth direction via SVD, revealing relational laws and cross-model convergence, validated on multiple models including Gemma-2-2b.
Francesco Karim Vicidomini
B2B uses EnergyPlus to generate 6,000 buildings and evaluates RL transfer across goals, dynamics, action spaces, and domains.
Vincent Taboga, Justin Veilleux, Doseok Jang et al.
Proposes Recursive Harness Self-Improvement (RHI), an iterative, lightweight method that significantly boosts low-reasoning agents' performance with minimal updates, reducing inference costs by up to 60%.
Hyunin Lee, Jinglue Xu, Jeffrey Seely et al.
Introduces multi-axis Max@K reinforcement learning to enhance target mode coverage in text-to-image models, significantly improving diversity and fairness.
Ku Onoda, Paavo Parmas, Hiroki Furuta et al.
Introduces a continuous-time RL framework to optimize discrete diffusion models, enhancing performance in reasoning and coding tasks.
Zikun Zhang, Jiayuan Sheng, David D. Yao et al.
TRACE assigns turn-level rewards via credit estimation using frozen reference models and TD changes, boosting long-horizon agent performance without extra critic training.
Leitian Tao, Baolin Peng, Wenlin Yao et al.
Introduced Self-Correcting Coupled Markov Jump Processes for concurrent image understanding and generation.
Minh-Quan Le, Armand Comas, Alexandros Lattas et al.
CoDiffGRN employs a co-evolutionary discrete diffusion model with BEELINE-KGC benchmark, significantly improving novel gene regulatory prediction.
Jiaze Song, Runhao Zhao, Minghao Xu et al.
SinAE: A single-architecture flow-matching autoencoder significantly reduces reconstruction errors across atomic systems.
Yuxuan Ren, Fan Yang, Jianhua Yao et al.
This paper introduces SOAP and Muon optimizers, with algorithmic improvements enabling stable, efficient large-scale LLM pretraining at billion-parameter scales.
Mikail Khona, Aditya Vavre, Boxiang Wang et al.
Quantized LLM reliability varies with bitwidth; 4-bit models offer the best efficiency-reliability trade-off.
Sirine Ayadi, Sándor Daróczi, Stephan Günnemann et al.
This paper improves convergence rates of single-loop AID and ITD for bilevel optimization from O(κ^6/K) to O(κ^5/K) and error from κ^3 to κ^2, using decoupled norm analysis.
Yubo Zhou, Jun Shu, Luo Luo et al.
Mach-Mind-4-Flash, a 35B MoE model with 3B activated parameters, matches 100B-class performance.
Foundation Model Team
Proposes risk-aware GUMDP framework using ERM and MCTS, enabling robust multi-task decision-making under uncertainty.
Pedro P. Santos, Fábio Vital, Alberto Sardinha et al.
Super method uses Wanda score for sparse fine-tuning, excelling on Math17K dataset.
Ivan Ilin, Philip Zmushko, Peter Richtárik
Proposes CABS-C and CABS-D algorithms leveraging surrogate rewards and correlations to improve LLM routing efficiency and reduce regret.
Ajay Narayanan Sridhar, Ronak Singh, Mehrdad Mahdavi et al.
Proposes SATS and Token Routing to improve activation sparsification in LLMs, enhancing efficiency and performance.
Bishmoy Paul, Youngmin Yi, Hoeseok Yang