$f$-Divergence Regularized RLHF: Two Tales of Sampling and Unified Analyses
Abstract
Lay Summary
Large language models, such as chatbots, can be improved by learning from people’s preferences between different answers. While improving the model, we also want to prevent it from moving too far away from its original behavior in unexpected or harmful ways. Most existing theory studies the common way to control this change, while recent practical methods have started using other ways that may be more stable or robust. This paper develops a general framework for understanding many such control methods, instead of analyzing each one separately. It proposes two strategies for deciding what feedback the model should collect while it is learning: one focuses on places where the model is uncertain, and the other uses how sensitive the learning process is to small changes as a guide. The paper proves that both strategies can learn efficiently. Overall, this work helps explain how different ways of controlling human-feedback training work, and provides researchers with principled tools for designing more reliable AI systems.