Reinforcement Learning without Ground-Truth Solutions can Improve LLMs
RiVER leverages score-based optimization with instance ranking and winner-heavy rewards, improving LLMs on algorithmic tasks by 8.9%-9.4%.
Yingyu Lin, Qiyue Gao, Nikki Lijing Kuang et al.