Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
This survey formalizes RLHF’s three-stage pipeline and shows why human feedback is not a sufficient safety guarantee.
Stephen Casper, Xander Davies, Claudia Shi et al.