Research map / Bandits and online learning
RL theory and MDPs: research map
271 accepted papers on RL theory and MDPs in Bandits and online learning, from ICML, NeurIPS, ICLR, CVPR and AAAI (2016–2026), grouped into 2 clusters and 6 approaches. The busiest year so far is 2021.
Within Bandits and online learning, its share shrank from 13.1% in 2023–24 to 10.1% in 2025–26 (73 → 60 papers at ICML, NeurIPS, CVPR and AAAI, the venues with data for all four years).
Explore RL theory and MDPs in the interactive map
Working on something in this topic? Describe your idea in scime atlas to see which approach it falls under, the closest papers by meaning and how crowded the spot has become.
Approaches and key papers
rl · function approximation · linear function · 154 papers
Approaches in this cluster:
- Regret bounds with function approximation (79 papers)
Prove instance-dependent and low-switching regret bounds for RL with function approximation. - Linear function approximation in RL (64 papers)
Establish efficient optimistic algorithms for RL with linear and generalized linear features. - Stochastic shortest path (11 papers)
Derive minimax-optimal regret for stochastic shortest path problems.
Most cited and most cited since 2024:
- Unifying PAC and Regret: Uniform PAC Bounds for Episodic Reinforcement Learning (NeurIPS 2017 · 138 citations)
- Model-Based Reinforcement Learning with Value-Targeted Regression (ICML 2020 · 63 citations)
- Dynamic Regret of Adversarial MDPs with Unknown Transition and Linear Function Approximation (AAAI 2024 · 1 citations)
- Provably Efficient Exploration in Quantum Reinforcement Learning with Logarithmic Worst-Case Regret (ICML 2024 · 1 citations)
markov decision · decision processes · processes mdps · 117 papers
Approaches in this cluster:
- Average-reward and latent MDPs (57 papers)
Derive regret guarantees for constrained, average-reward, factored, and latent MDPs. - Non-stationary and structured MDPs (30 papers)
Learn unknown, non-stationary, or structured MDPs with provable efficiency. - Adversarial MDPs with bandit feedback (30 papers)
Achieve near-optimal regret in adversarial and delayed-feedback MDPs.
Most cited and most cited since 2024:
- Near Optimal Exploration-Exploitation in Non-Communicating Markov Decision Processes (NeurIPS 2018 · 58 citations)
- Learning Unknown Markov Decision Processes: A Thompson Sampling Approach (NeurIPS 2017 · 54 citations)
- Regret Analysis of Policy Gradient Algorithm for Infinite Horizon Average Reward Markov Decision Processes (AAAI 2024 · 2 citations)
- Achieving Tractable Minimax Optimal Regret in Average Reward MDPs (NeurIPS 2024 · 2 citations)
