Training language models to follow instructions with human feedback
Long Ouyang, Ryan Lowe(Australian Research Council), Peter Welinder(California Institute of Technology), Fraser Kelton, Maddie Simens, Carroll L. Wainwright, Katarina Slama, Pamela Mishkin, Amanda Askell, Luke E. Miller(Lyon 1 Université), Sandhini Agarwal(OpenAI (United States)), Alex Ray, Jan Leike, Chong Zhang, John Schulman, Xu Jiang(Chinese Academy of Agricultural Sciences), Diogo Almeida, Jeff Wu, Paul Christiano, Jacob Hilton(National Institutes of Health)
Cited by 4,321
Related Papers
Global warming and recurrent mass bleaching of corals
|Nature|2017|3.4k
Learning dexterous in-hand manipulation
|The International Journal of Robotics Research|2019|1.6k
Evaluating Large Language Models Trained on Code
|arXiv (Cornell University)|2021|1.5k
Training Language Models to Follow Instructions with Human Feedback
|Unknown|2022|760