Research map / Bandits and online learning
Contextual bandits and Thompson sampling: research map
546 accepted papers on Contextual bandits and Thompson sampling in Bandits and online learning, from ICML, NeurIPS, ICLR, CVPR and AAAI (2016–2026), grouped into 2 clusters and 7 approaches. The busiest year so far is 2024.
Within Bandits and online learning, its share shrank from 26.8% in 2023–24 to 22.4% in 2025–26 (149 → 133 papers at ICML, NeurIPS, CVPR and AAAI, the venues with data for all four years).
Explore Contextual bandits and Thompson sampling in the interactive map
Working on something in this topic? Describe your idea in scime atlas to see which approach it falls under, the closest papers by meaning and how crowded the spot has become.
Approaches and key papers
contextual bandits · contexts · user · 436 papers
Approaches in this cluster:
- Oracle-efficient contextual bandits (168 papers)
Build practical contextual bandit algorithms using regression oracles and handling misspecification or privacy. - Linear and generalized linear bandits (88 papers)
Derive improved regret bounds for stochastic linear bandits with efficient updates. - Practical contextual bandit extensions (77 papers)
Tune, infer from, cluster, or meta-learn contextual bandit algorithms. - Causal and structured bandits (61 papers)
Exploit causal graphs and auxiliary feedback to learn good interventions. - Ranking bandits (42 papers)
Learn to rank online from click feedback with unimodal and semi-bandit methods.
Most cited and most cited since 2024:
- Fairness in Learning: Classic and Contextual Bandits (NeurIPS 2016 · 191 citations)
- Provably Optimal Algorithms for Generalized Linear Contextual Bandits (ICML 2017 · 144 citations)
- Using Adaptive Bandit Experiments to Increase and Investigate Engagement in Mental Health (AAAI 2024 · 11 citations)
- Causal Bandits for Linear Structural Equation Models (NeurIPS 2024 · 5 citations)
thompson sampling · bayesian · ts · 110 papers
Approaches in this cluster:
- Thompson sampling analysis (84 papers)
Analyze and improve Thompson sampling for combinatorial and high-dimensional contextual bandits. - Gaussian process bandit regret (26 papers)
Tighten regret bounds for GP-UCB and Thompson sampling in Gaussian process optimization.
Most cited and most cited since 2024:
- Gaussian Process Bandit Optimisation with Multi-fidelity Evaluations (NeurIPS 2016 · 63 citations)
- Improved Regret Bounds for Thompson Sampling in Linear Quadratic Control Problems (ICML 2018 · 52 citations)
- Learning Encodings for Constructive Neural Combinatorial Optimization Needs to Regret (AAAI 2024 · 6 citations)
- The Choice of Noninformative Priors for Thompson Sampling in Multiparameter Bandit Models (AAAI 2024 · 2 citations)
