atlas

Research map / Bandits and online learning

RL theory and MDPs: research map

271 accepted papers on RL theory and MDPs in Bandits and online learning, from ICML, NeurIPS, ICLR, CVPR and AAAI (2016–2026), grouped into 2 clusters and 6 approaches. The busiest year so far is 2021.

Within Bandits and online learning, its share shrank from 13.1% in 2023–24 to 10.1% in 2025–26 (73 → 60 papers at ICML, NeurIPS, CVPR and AAAI, the venues with data for all four years).

2016: 1162017: 9172018: 4182019: 14192020: 23202021: 46212022: 41222023: 37232024: 36242025: 32252026: 2826

Explore RL theory and MDPs in the interactive map

Working on something in this topic? Describe your idea in scime atlas to see which approach it falls under, the closest papers by meaning and how crowded the spot has become.

Approaches and key papers

rl · function approximation · linear function · 154 papers

Approaches in this cluster:

Most cited and most cited since 2024:

markov decision · decision processes · processes mdps · 117 papers

Approaches in this cluster:

Most cited and most cited since 2024:

Related topics in Bandits and online learning