Mitigating Reward Hacking in RLHF via Bayesian Non-negative Reward Modeling
BNRM integrates Bayesian non-negative factor analysis to mitigate reward hacking in RLHF, improving robustness and interpretability.
Zhibin Duan, Guowei Rong, Zhuo Li et al.
BNRM integrates Bayesian non-negative factor analysis to mitigate reward hacking in RLHF, improving robustness and interpretability.
Zhibin Duan, Guowei Rong, Zhuo Li et al.
DAWN method improves value learning efficiency in residual RL through data anchoring and normalization.
Guozheng Ma, Lu Li, Haoyu Wang et al.
RLTT distributes rewards across latent reasoning steps, significantly boosting LoopLM's performance on math and reasoning benchmarks.
Jonathan Williams, Esin Tureci
WildCat uses randomly pivoted Cholesky to select a small weighted coreset for near-linear time attention.
Tobias Schröder, Lester Mackey
Proposes ExO-PPO, combining conservative on-policy guarantees with off-policy data reuse to enhance sample efficiency and stability.
Hanyong Wang, Menglong Yang
The Physical Information Bottleneck (PIB) framework enhances energy efficiency and speed in deep physical neural networks.
Hao Wang, Ziao Wang, Xiangpeng Liang et al.
AnomSeer combines ExpCoT and TimerPO to enhance fine-grained reasoning and detection accuracy in TSAD, outperforming baselines with 58.8% F1.
Junru Zhang, Lang Feng, Haoran Shi et al.
LEFT fuses tri-view features with analysis-synthesis cycle consistency for unsupervised time series anomaly detection, achieving over 3% ROC and 6% PR improvements.
Dezheng Wang, Tong Chen, Guansong Pang et al.
RLTR introduces cross-model transfer rewards to enhance reasoning robustness and efficiency, achieving +3.6% in Maj@64 on MATH-500 and 2.5× training step reduction.
Hyunseok Lee, Soheil Abbasloo, Jihoon Tack et al.
SkillRL enhances performance by 15.3% in tasks like ALFWorld via recursive skill-augmented reinforcement learning.
Peng Xia, Jianwen Chen, Hanyang Wang et al.
TAAM leverages lightweight Neural Synapse Modulators and Anchored Multi-hop Propagation for replay-free, resource-efficient continual graph learning, outperforming SOTA with zero forgetting.
Jingtao Liu, Xinming Zhang
OGPSA uses orthogonal gradient projection to mitigate capability loss during LLM safety alignment, improving safety-utility trade-off by over 9%.
Guanglong Sun, Siyuan Zhang, Liyuan Wang et al.
rePIRL learns process reward models via inverse RL to enhance LLM reasoning.
Xian Wu, Kaijie Zhu, Ying Zhang et al.
Proposes Learnable Chernoff Baselines (LCBs) for inference-time model alignment, reducing queries by 7x, with total variation guarantees.
Sunil Madhow, Yuchen Liang, Ness Shroff et al.
ShaPO enhances LLM safety alignment robustness via selective geometry control, outperforming benchmarks.
Yonghui Yang, Wenjian Tao, Jilong Liu et al.
Using thermodynamic entropy production, proposes EDS and WDS schedules, significantly improving discrete diffusion sampling efficiency.
Alberto Foresti, Mustapha Bounoua, Giulio Franzese et al.
SNIP, Wanda, SafeNeuron, and NLSR produced IoUs of 0.01–0.72, revealing no stable dataset-agnostic safety region.
Zongmin Li, Jian Su, Farah Benamara et al.
Proposes Soft FB, a maximum entropy variant for zero-shot RL, enabling direct optimization of general differentiable utilities via offline data and low-dimensional search.
Marco Bagatella, Thomas Rupf, Georg Martius et al.
Diamond Maps employs stochastic flow maps for efficient reward alignment, enabling rapid adaptation during inference with minimal computational overhead.
Peter Holderrieth, Douglas Chen, Luca Eyring et al.
A-GRAE dynamically adjusts exploration incentives and sample difficulty focus, improving GRPO's exploration and adaptation in complex tasks.
Zhiqi Yu, Zhangquan Chen, Mengting Liu et al.