Paper
Training language models to follow instructions with human feedback (InstructGPT, 2022): the RLHF recipe behind ChatGPT
Ouyang et al. (OpenAI) fine-tune GPT-3 on human demonstrations, train a reward model on human preferences, then optimize the policy with PPO. The 1.3B InstructGPT is preferred over the 175B GPT-3, and the three-step RLHF recipe became how every assistant model is aligned. 68 pages.