Research map / Bandits and online learning
Multi-armed bandits: research map
497 accepted papers on Multi-armed bandits in Bandits and online learning, from ICML, NeurIPS, ICLR, CVPR and AAAI (2016–2026), grouped into 2 clusters and 7 approaches. The busiest year so far is 2021.
Within Bandits and online learning, its share grew from 19.0% in 2023–24 to 21.0% in 2025–26 (106 → 125 papers at ICML, NeurIPS, CVPR and AAAI, the venues with data for all four years).
Explore Multi-armed bandits in the interactive map
Working on something in this topic? Describe your idea in scime atlas to see which approach it falls under, the closest papers by meaning and how crowded the spot has become.
Approaches and key papers
multi armed · armed bandit · stochastic · 348 papers
Approaches in this cluster:
- Non-stationary and structured-reward bandits (140 papers)
Handle rising, restless, blocking, and recovering reward dynamics in bandit algorithms. - Optimal stochastic bandit algorithms (124 papers)
Characterize regret-optimal algorithms for stochastic, batched, and combinatorial bandits. - Multi-agent bandits (42 papers)
Coordinate multiple or federated agents in decentralized bandit learning. - Bandit formulations for ML problems (42 papers)
Cast LLM decoding, sampling, and experimental design as bandit problems.
Most cited and most cited since 2024:
- Minimal Exploration in Structured Stochastic Bandits (NeurIPS 2017 · 79 citations)
- Combinatorial Multi-Armed Bandit with General Reward Functions (NeurIPS 2016 · 75 citations)
- Exploration, Exploitation, and Engagement in Multi-Armed Bandits with Abandonment (NeurIPS 2024 · 8 citations)
- Online Restless Multi-Armed Bandits with Long-Term Fairness Constraints (AAAI 2024 · 7 citations)
identification · best arm · pure exploration · 149 papers
Approaches in this cluster:
- Best-arm identification in linear bandits (57 papers)
Identify the best arm with optimal sample complexity for linear and fixed-budget settings. - Pure exploration algorithms (57 papers)
Design efficient pure-exploration and thresholding algorithms with batch and combinatorial structure. - Fixed-confidence versus fixed-budget identification (35 papers)
Develop optimal elimination and top-two algorithms for confidence and budget settings.
Most cited and most cited since 2024:
- An optimal algorithm for the Thresholding Bandit Problem (ICML 2016 · 64 citations)
- Refined Lower Bounds for Adversarial Bandits (NeurIPS 2016 · 55 citations)
- Thompson Sampling for Real-Valued Combinatorial Pure Exploration of Multi-Armed Bandit (AAAI 2024 · 5 citations)
- Quantum Best Arm Identification with Quantum Oracles (AAAI 2025 · 2 citations)
