Optimizers: research map
1,124 accepted papers on Optimizers in Optimization, from ICML, NeurIPS, ICLR, CVPR and AAAI (2016–2026), grouped into 4 clusters and 16 approaches. The busiest year so far is 2025.
Within Optimization, its share grew from 27.9% in 2023–24 to 29.7% in 2025–26 (242 → 290 papers at ICML, NeurIPS, CVPR and AAAI, the venues with data for all four years).
Explore Optimizers in the interactive map
Working on something in this topic? Describe your idea in scime atlas to see which approach it falls under, the closest papers by meaning and how crowded the spot has become.
Approaches and key papers
differentiable · functions · loss · 438 papers
Approaches in this cluster:
- Differentiating through embedded optimization (117 papers)
Backpropagates through optimization layers and learns optimizers or proximal operators. - Surrogate losses and robust solvers (119 papers)
Proposes tailored losses, quantization regularizers, and fast solvers for specific learning problems. - Constrained and non-backprop training (97 papers)
Trains networks with orthogonality constraints, growth steps, or alternating minimization instead of plain backpropagation. - Randomized gradient estimation (77 papers)
Reduces variance in gradient estimators via randomized differentiation, evolution strategies, and weight averaging. - Multi-objective gradient methods (28 papers)
Provides stochastic gradient algorithms with convergence guarantees for multi-objective and multi-task learning.
Most cited and most cited since 2024:
- GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium (NeurIPS 2017 · 4,380 citations)
- Deep Leakage from Gradients (NeurIPS 2019 · 1,528 citations)
- Runtime Analysis of the SMS-EMOA for Many-Objective Optimization (AAAI 2024 · 38 citations)
- LP++: A Surprisingly Strong Linear Probe for Few-Shot CLIP (CVPR 2024 · 30 citations)
adaptive · optimizers · momentum · 299 papers
Approaches in this cluster:
- Adam convergence and adaptive learning rates (78 papers)
Analyzes Adam-family convergence and adapts learning rates using hypergradients. - Memory-efficient LLM optimizers (79 papers)
Designs large-batch, low-memory, and normalized optimizers for training large language models. - SGD momentum and step-size schedules (72 papers)
Studies momentum, batch size, and step-size schedules for stochastic gradient descent. - Second-order and structured preconditioning (70 papers)
Builds diagonal, full-matrix, and Hessian-based preconditioners for adaptive optimization.
Most cited and most cited since 2024:
- Decoupled Weight Decay Regularization (ICLR 2019 · 9,419 citations)
- On the Convergence of Adam and Beyond (ICLR 2018 · 1,616 citations)
- Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training (ICLR 2024 · 28 citations)
- The Road Less Scheduled (NeurIPS 2024 · 16 citations)
variational · sampling · inference · 289 papers
Approaches in this cluster:
- Stochastic gradient MCMC (142 papers)
Uses stochastic gradient Langevin dynamics and related samplers for scalable Bayesian posterior sampling. - Natural gradient and Fisher approximations (77 papers)
Develops and critiques Kronecker-factored and structured approximations of natural gradient and second-order updates. - Stein variational gradient descent (47 papers)
Transports particles with kernelized Stein variational updates for Bayesian inference, with analysis and extensions. - Zeroth-order training methods (23 papers)
Trains models using gradient estimates from function queries, with convergence and learning-dynamics analysis.
Most cited and most cited since 2024:
- Stein Variational Gradient Descent: A General Purpose Bayesian Inference Algorithm (NeurIPS 2016 · 307 citations)
- A Variational Inequality Perspective on Generative Adversarial Networks (ICLR 2019 · 157 citations)
- PseuZO: Pseudo-Zeroth-Order Algorithm for Training Deep Neural Networks (NeurIPS 2025 · 3 citations)
- On Estimating the Gradient of the Expected Information Gain in Bayesian Experimental Design (AAAI 2024 · 3 citations)
bilevel · inner · shot · 98 papers
Approaches in this cluster:
- Gradient-based meta-learning theory (40 papers)
Provides guarantees and improved algorithms for MAML-style outer-loop meta-learning. - Hypergradient bilevel optimization (35 papers)
Computes approximate hypergradients to solve bilevel problems for hyperparameter optimization. - Differentiating through inner loops (23 papers)
Optimizes hyperparameters over long horizons using implicit differentiation and truncated inner-loop gradients.
Most cited and most cited since 2024:
- Optimization as a Model for Few-Shot Learning (ICLR 2017 · 2,375 citations)
- Learning to Reweight Examples for Robust Deep Learning (ICML 2018 · 841 citations)
- MetaDiff: Meta-Learning with Conditional Diffusion for Few-Shot Learning (AAAI 2024 · 68 citations)
- Towards Consistent Multi-Task Learning: Unlocking the Potential of Task-Specific Parameters (CVPR 2025 · 7 citations)
