Bandits and online learning: research map
Regret-minimizing decisions under uncertainty, from multi-armed bandits to RL theory.
2,266 accepted papers at ICML, NeurIPS, ICLR, CVPR and AAAI (2016–2026), in 4 topics. Within all five venues, its share held steady from 2.1% in 2023–24 to 1.3% in 2025–26 (557 → 594 papers at ICML, NeurIPS, CVPR and AAAI, the venues with data for all four years).
Explore Bandits and online learning in the interactive map
Topics
- Online convex optimization and games · 952 papers
Online learning with feedback graphs, Best-of-both-worlds bandit algorithms, Learning for pricing and auctions - Contextual bandits and Thompson sampling · 546 papers
Oracle-efficient contextual bandits, Linear and generalized linear bandits, Practical contextual bandit extensions - Multi-armed bandits · 497 papers
Non-stationary and structured-reward bandits, Optimal stochastic bandit algorithms, Multi-agent bandits - RL theory and MDPs · 271 papers
Regret bounds with function approximation, Linear function approximation in RL, Stochastic shortest path
Most cited papers in Bandits and online learning
- Fairness in Learning: Classic and Contextual Bandits (NeurIPS 2016 · 191 citations)
- Provably Optimal Algorithms for Generalized Linear Contextual Bandits (ICML 2017 · 144 citations)
- Unifying PAC and Regret: Uniform PAC Bounds for Episodic Reinforcement Learning (NeurIPS 2017 · 138 citations)
- Online Meta-Learning (ICML 2019 · 93 citations)
- Adaptive Algorithms for Online Convex Optimization with Long-term Constraints (ICML 2016 · 82 citations)
- Minimal Exploration in Structured Stochastic Bandits (NeurIPS 2017 · 79 citations)
- Combinatorial Multi-Armed Bandit with General Reward Functions (NeurIPS 2016 · 75 citations)
- Online convex optimization for cumulative constraints (NeurIPS 2018 · 72 citations)
- Dynamic Incentive-Aware Learning: Robust Pricing in Contextual Auctions (NeurIPS 2019 · 71 citations)
- An optimal algorithm for the Thresholding Bandit Problem (ICML 2016 · 64 citations)
- Gaussian Process Bandit Optimisation with Multi-fidelity Evaluations (NeurIPS 2016 · 63 citations)
- Model-Based Reinforcement Learning with Value-Targeted Regression (ICML 2020 · 63 citations)
