Training language models to follow instructions with human feedback

Long Ouyang, Ryan Lowe(Australian Research Council), Peter Welinder(California Institute of Technology), Fraser Kelton, Maddie Simens, Carroll L. Wainwright, Katarina Slama, Pamela Mishkin, Amanda Askell, Luke E. Miller(Lyon 1 Université), Sandhini Agarwal(OpenAI (United States)), Alex Ray, Jan Leike, Chong Zhang, John Schulman, Xu Jiang(Chinese Academy of Agricultural Sciences), Diogo Almeida, Jeff Wu, Paul Christiano, Jacob Hilton(National Institutes of Health)
arXiv (Cornell University)
March 4, 2022
Cited by 4,321


Related Papers