Human-level control through deep reinforcement learningVolodymyr Mnih, Demis Hassabis, Ioannis Antonoglou et al.|Nature|2015Cited by 29.9k
Playing Atari with Deep Reinforcement LearningVolodymyr Mnih, Martin Riedmiller, Koray Kavukcuoglu et al.|arXiv (Cornell University)|2013Cited by 5.1k
Deterministic policy gradient algorithmsDavid Silver, Martin Riedmiller, Guy Lever et al.|HAL (Le Centre pour la Communication Scientifique Directe)|2014Cited by 1.7k
Solving Deep Memory POMDPs with Recurrent Policy GradientsDaan Wierstra, Jürgen Schmidhuber, A. Foerster et al.|Lecture notes in computer science|2007Cited by 149