Implicit bias and generalization: research map
958 accepted papers on Implicit bias and generalization in Optimization, from ICML, NeurIPS, ICLR, CVPR and AAAI (2016–2026), grouped into 3 clusters and 12 approaches. The busiest year so far is 2026.
Within Optimization, its share held steady from 27.1% in 2023–24 to 26.7% in 2025–26 (235 → 261 papers at ICML, NeurIPS, CVPR and AAAI, the venues with data for all four years).
Explore Implicit bias and generalization in the interactive map
Working on something in this topic? Describe your idea in scime atlas to see which approach it falls under, the closest papers by meaning and how crowded the spot has become.
Approaches and key papers
relu · width · kernel · 434 papers
Approaches in this cluster:
- Shallow ReLU network convergence (132 papers)
Proves gradient descent convergence and generalization for shallow and deep ReLU or linear networks. - Neural tangent kernel analysis (114 papers)
Uses NTK and linearization theory to analyze training, feature learning, and inductive bias in wide networks. - Gradient flow and mean-field dynamics (77 papers)
Studies training dynamics of wide networks via gradient flow, mean-field limits, and teacher-student models. - Low-rank and single-index recovery (70 papers)
Analyzes gradient descent on matrix sensing, factorization, and single-index models with global guarantees. - In-context learning as gradient descent (41 papers)
Shows transformers implement gradient-descent-like algorithms for in-context linear regression.
Most cited and most cited since 2024:
- Neural Tangent Kernel: Convergence and Generalization in Neural Networks (NeurIPS 2018 · 1,345 citations)
- A Convergence Theory for Deep Learning via Over-Parameterization (ICML 2019 · 688 citations)
- How does PDE order affect the convergence of PINNs? (NeurIPS 2024 · 5 citations)
- Neural Network-Based Score Estimation in Diffusion Models: Optimization and Generalization (ICLR 2024 · 3 citations)
sgd · sharpness · noise · 353 papers
Approaches in this cluster:
- SGD noise and flat minima (112 papers)
Analyzes how stochastic gradient noise structure and learning rate steer SGD toward flat minima. - Edge-of-stability dynamics (116 papers)
Characterizes gradient descent at large step sizes where sharpness hovers at the stability threshold. - Stability-based generalization analysis (88 papers)
Bounds generalization of gradient-trained networks through algorithmic stability and Bayesian arguments. - Sharpness-aware minimization theory (37 papers)
Analyzes why and how sharpness-aware minimization works and proposes improved variants.
Most cited and most cited since 2024:
- Understanding deep learning requires rethinking generalization (ICLR 2017 · 1,040 citations)
- Visualizing the Loss Landscape of Neural Nets (NeurIPS 2018 · 606 citations)
- Friendly Sharpness-Aware Minimization (CVPR 2024 · 22 citations)
- Why Warmup the Learning Rate? Underlying Mechanisms and Improvements (NeurIPS 2024 · 15 citations)
implicit bias · regularization · bias gradient · 171 papers
Approaches in this cluster:
- Explicit and implicit regularization interplay (63 papers)
Studies how weight decay, gradient regularizers, and training dynamics jointly induce implicit regularization. - Implicit regularization of SGD (55 papers)
Analyzes implicit regularization from SGD in least squares, matrix factorization, and diagonal networks. - Implicit bias in linear networks (53 papers)
Characterizes the max-margin-like solutions gradient flow selects in linear and separable classification settings.
Most cited and most cited since 2024:
- Characterizing Implicit Bias in Terms of Optimization Geometry (ICML 2018 · 168 citations)
- On the Optimization of Deep Networks: Implicit Acceleration by Overparameterization (ICML 2018 · 136 citations)
- Why Do We Need Weight Decay in Modern Deep Learning? (NeurIPS 2024 · 10 citations)
- Controlling the Flow: Stability and Convergence for Stochastic Gradient Descent with Decaying Regularization (NeurIPS 2025 · 2 citations)
Related topics in Optimization
- Distributed and compressed (349)
- Nonconvex and smooth optimization (1,448)
- Optimizers (1,124)
