Research map / Bandits and online learning
Online convex optimization and games: research map
952 accepted papers on Online convex optimization and games in Bandits and online learning, from ICML, NeurIPS, ICLR, CVPR and AAAI (2016–2026), grouped into 3 clusters and 13 approaches. The busiest year so far is 2026.
Within Bandits and online learning, its share grew from 41.1% in 2023–24 to 46.5% in 2025–26 (229 → 276 papers at ICML, NeurIPS, CVPR and AAAI, the venues with data for all four years).
Explore Online convex optimization and games in the interactive map
Working on something in this topic? Describe your idea in scime atlas to see which approach it falls under, the closest papers by meaning and how crowded the spot has become.
Approaches and key papers
price · pricing · auctions · 648 papers
Approaches in this cluster:
- Online learning with feedback graphs (155 papers)
Design best-of-both-worlds experts and bandit algorithms with structured feedback and limited information. - Best-of-both-worlds bandit algorithms (125 papers)
Obtain optimal regret in both stochastic and adversarial regimes for linear and semi-bandits. - Learning for pricing and auctions (101 papers)
Develop regret-minimizing pricing and bidding algorithms for repeated auctions with strategic buyers. - Online allocation with constraints (77 papers)
Solve online allocation and knapsack problems using dual mirror descent and primal-dual methods. - Online decision-making with feedback structure (81 papers)
Analyze online prediction under delayed, bandit, strategic, or dynamical feedback. - Online linear-quadratic control regret (36 papers)
Learn linear-quadratic regulators online with regret guarantees. - Learning-augmented online algorithms (73 papers)
Use machine-learned predictions to improve online algorithms while preserving worst-case robustness.
Most cited and most cited since 2024:
- Online Meta-Learning (ICML 2019 · 93 citations)
- Dynamic Incentive-Aware Learning: Robust Pricing in Contextual Auctions (NeurIPS 2019 · 71 citations)
- Ahpatron: A New Budgeted Online Kernel Learning Machine with Tighter Mistake Bound (AAAI 2024 · 4 citations)
- Dynamic Budget Throttling in Repeated Second-Price Auctions (AAAI 2024 · 4 citations)
convex optimization · online convex · functions · 187 papers
Approaches in this cluster:
- Constrained and unconstrained OCO algorithms (85 papers)
Design online convex optimization algorithms handling constraints, unbounded domains, and switching costs. - Dynamic and adaptive regret (57 papers)
Reduce and bound dynamic and adaptive regret for convex and smooth losses. - Parameter-free non-stationary online learning (45 papers)
Achieve universal, parameter-free guarantees in changing environments through meta-algorithm ensembles.
Most cited and most cited since 2024:
- Adaptive Algorithms for Online Convex Optimization with Long-term Constraints (ICML 2016 · 82 citations)
- Online convex optimization for cumulative constraints (NeurIPS 2018 · 72 citations)
- Optimal Algorithms for Online Convex Optimization with Adversarial Constraints (NeurIPS 2024 · 7 citations)
- Online Submodular Maximization via Online Convex Optimization (AAAI 2024 · 5 citations)
game · equilibrium · convergence · 117 papers
Approaches in this cluster:
- Last-iterate convergence in games (65 papers)
Prove convergence of uncoupled no-regret dynamics in zero-sum, monotone, and potential games. - Fast no-regret dynamics in games (37 papers)
Obtain near-constant regret and fast convergence via optimism and swap-regret analysis. - Counterfactual regret minimization (15 papers)
Improve CFR for extensive-form games with neural approximation and learned algorithm design.
Most cited and most cited since 2024:
- Learning with Bandit Feedback in Potential Games (NeurIPS 2017 · 53 citations)
- Learning in Games: Robustness of Fast Convergence (NeurIPS 2016 · 44 citations)
- Learning Not to Regret (AAAI 2024 · 2 citations)
- Optimism Without Regularization: Constant Regret in Zero-Sum Games (NeurIPS 2025 · 1 citations)
