Reinforcement learning: research map
Policy learning, offline RL, imitation and learning from feedback.
6,108 accepted papers at ICML, NeurIPS, ICLR, CVPR and AAAI (2016–2026), in 6 topics. Within all five venues, its share held steady from 5.9% in 2023–24 to 5.1% in 2025–26 (1,568 → 2,317 papers at ICML, NeurIPS, CVPR and AAAI, the venues with data for all four years).
Explore Reinforcement learning in the interactive map
Topics
- Visual and representation RL · 2,133 papers
In-context and adaptive RL, Scalable RL systems and environments, Probabilistic and autonomous RL - Policy gradient and actor-critic · 1,502 papers
Model-based RL with guarantees, Regularized scalable deep value learning, Temporal-difference and discounting theory - RL from feedback · 737 papers
RL for LLM reasoning, Scaling RL with verifiable rewards, GRPO stability and variants - MDPs and planning · 682 papers
Average-reward and regularized MDPs, POMDPs and confounded RL, Robust MDP policy gradient - Offline RL · 623 papers
Offline RL sample complexity theory, Simple practical offline RL, Pessimism and uncertainty in offline RL - Imitation learning · 431 papers
Offline imitation from imperfect data, Imitation combined with RL, Behavior cloning theory and models
Most cited papers in Reinforcement learning
- Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor (ICML 2018 · 3,754 citations)
- Addressing Function Approximation Error in Actor-Critic Methods (ICML 2018 · 2,316 citations)
- Self-Critical Sequence Training for Image Captioning (CVPR 2017 · 2,268 citations)
- Dueling Network Architectures for Deep Reinforcement Learning (ICML 2016 · 1,777 citations)
- A Deep Reinforced Model for Abstractive Summarization (ICLR 2018 · 1,255 citations)
- Benchmarking Deep Reinforcement Learning for Continuous Control (ICML 2016 · 949 citations)
- Hierarchical Deep Reinforcement Learning: Integrating Temporal Abstraction and Intrinsic Motivation (NeurIPS 2016 · 822 citations)
- Guided Cost Learning: Deep Inverse Optimal Control via Policy Optimization (ICML 2016 · 591 citations)
- Action-Decision Networks for Visual Tracking With Deep Reinforcement Learning (CVPR 2017 · 574 citations)
- Conservative Q-Learning for Offline Reinforcement Learning (NeurIPS 2020 · 556 citations)
- Deep Reinforcement Learning from Human Preferences (NeurIPS 2017 · 500 citations)
- Decision Transformer: Reinforcement Learning via Sequence Modeling (NeurIPS 2021 · 464 citations)
