Training Language Models to Follow Instructions with Human Feedback
Long Ouyang, Ryan Lowe, Maddie Simens, Amanda Askell, Paul F. Christiano, Sandhini Agarwal(OpenAI (United States)), Diogo Almeida(Universidade Federal de Pernambuco), Katarina Slama, Alex Ray, Luke Miller, Pamela Mishkin, John Schulman, Chong Zhang, Jacob Hilton, Jeffrey Wu, Jan Leike, Carroll Wainwright, Peter Welinder(California Institute of Technology), Xu Jiang, Fraser Kelton
Cited by 760
Related Papers
Training language models to follow instructions with human feedback
|arXiv (Cornell University)|2022|4.3k
Learning dexterous in-hand manipulation
|The International Journal of Robotics Research|2019|1.6k
The Multidimensional Wisdom of Crowds
|CaltechAUTHORS (California Institute of Technology)|2010|735
Solving Rubik's Cube with a Robot Hand
|arXiv (Cornell University)|2019|634
Cascaded pose regression
|Unknown|2010|554