Research map / Reinforcement learning
MDPs and planning: research map
682 accepted papers on MDPs and planning in Reinforcement learning, from ICML, NeurIPS, ICLR, CVPR and AAAI (2016–2026), grouped into 2 clusters and 8 approaches. The busiest year so far is 2024.
Within Reinforcement learning, its share shrank from 12.9% in 2023–24 to 7.6% in 2025–26 (203 → 175 papers at ICML, NeurIPS, CVPR and AAAI, the venues with data for all four years).
Explore MDPs and planning in the interactive map
Working on something in this topic? Describe your idea in scime atlas to see which approach it falls under, the closest papers by meaning and how crowded the spot has become.
Approaches and key papers
markov decision · processes · observable · 389 papers
Approaches in this cluster:
- Average-reward and regularized MDPs (135 papers)
Develop theory for regularized, average-reward, and non-stationary MDPs. - POMDPs and confounded RL (99 papers)
Study RL under partial observability and confounding. - Robust MDP policy gradient (87 papers)
Optimize policies in robust MDPs with rectangular uncertainty. - POMDP policy optimization (68 papers)
Solve POMDPs with Monte Carlo, transformers, and policy-based methods.
Most cited and most cited since 2024:
- Mitigating Bias in Face Recognition Using Skewness-Aware Reinforcement Learning (CVPR 2020 · 229 citations)
- A Theory of Regularized Markov Decision Processes (ICML 2019 · 116 citations)
- Learning Logic Specifications for Policy Guidance in POMDPs: an Inductive Logic Programming Approach (AAAI 2025 · 5 citations)
- Principal-Agent Reward Shaping in MDPs (AAAI 2024 · 5 citations)
sample complexity · function approximation · epsilon · 293 papers
Approaches in this cluster:
- Function approximation in RL theory (93 papers)
Establish sample-efficient RL with general and low-rank function approximation. - Offline RL sample complexity (86 papers)
Prove near-optimal offline RL guarantees with pessimism and variance reduction. - Policy optimization theory (61 papers)
Analyze policy finetuning, reward-free exploration, and effective horizon. - Rich-observation exploration (53 papers)
Explore provably with state abstraction in rich-observation and model-based RL.
Most cited and most cited since 2024:
- Is Q-Learning Provably Efficient? (NeurIPS 2018 · 306 citations)
- Contextual Decision Processes with low Bellman rank are PAC-Learnable (ICML 2017 · 143 citations)
- Overcoming the Sim-to-Real Gap: Leveraging Simulation to Learn to Explore for Real-World RL (NeurIPS 2024 · 11 citations)
- Provable and Practical: Efficient Exploration in Reinforcement Learning via Langevin Monte Carlo (ICLR 2024 · 5 citations)
Related topics in Reinforcement learning
- Visual and representation RL (2,133)
- Policy gradient and actor-critic (1,502)
- Imitation learning (431)
- RL from feedback (737)
- Offline RL (623)
