Research map / Reinforcement learning
Policy gradient and actor-critic: research map
1,502 accepted papers on Policy gradient and actor-critic in Reinforcement learning, from ICML, NeurIPS, ICLR, CVPR and AAAI (2016–2026), grouped into 5 clusters and 18 approaches. The busiest year so far is 2026.
Within Reinforcement learning, its share shrank from 21.9% in 2023–24 to 20.0% in 2025–26 (343 → 463 papers at ICML, NeurIPS, CVPR and AAAI, the venues with data for all four years).
Explore Policy gradient and actor-critic in the interactive map
Working on something in this topic? Describe your idea in scime atlas to see which approach it falls under, the closest papers by meaning and how crowded the spot has become.
Approaches and key papers
value · td · distributional · 773 papers
Approaches in this cluster:
- Model-based RL with guarantees (245 papers)
Analyze and design model-based RL algorithms with provable robustness and sample efficiency. - Regularized scalable deep value learning (216 papers)
Stabilize deep value-based RL through regularization, overfitting control, and scaling. - Temporal-difference and discounting theory (143 papers)
Analyze TD learning, discount factors, and value estimation error for prediction and control. - Tree search and policy optimization (109 papers)
Connect Monte-Carlo tree search and sequential Monte Carlo to regularized policy optimization. - Distributional value learning (60 papers)
Learn full return distributions with categorical, moment-matching, or spline-based representations.
Most cited and most cited since 2024:
- Dueling Network Architectures for Deep Reinforcement Learning (ICML 2016 · 1,777 citations)
- Guided Cost Learning: Deep Inverse Optimal Control via Policy Optimization (ICML 2016 · 591 citations)
- Diffusion-based Reinforcement Learning via Q-weighted Variational Policy Optimization (NeurIPS 2024 · 17 citations)
- Fractional Deep Reinforcement Learning for Age-Minimal Mobile Edge Computing (AAAI 2024 · 10 citations)
policy gradient · gradients · variance · 242 papers
Approaches in this cluster:
- Variance-reduced policy gradients (79 papers)
Reduce gradient variance with off-policy estimators, factorized baselines, and risk-sensitive formulations. - Trajectory reuse in policy gradients (68 papers)
Reuse past trajectories and aggregate gradients for faster, more stable policy optimization. - Alternative policy update rules (50 papers)
Replace standard gradient steps with evolved, natural-gradient, or divide-and-conquer updates. - Policy gradients for special settings (45 papers)
Extend policy gradients to continuous time, performative, bounded-action, and diffusion-policy settings.
Most cited and most cited since 2024:
- Self-Critical Sequence Training for Image Captioning (CVPR 2017 · 2,268 citations)
- Distributed Distributional Deterministic Policy Gradients (ICLR 2018 · 320 citations)
- Optimistic Policy Gradient in Multi-Player Markov Games with a Single Controller: Convergence beyond the Minty Property (AAAI 2024 · 4 citations)
- Excluding the Irrelevant: Focusing Reinforcement Learning through Continuous Action Masking (NeurIPS 2024 · 4 citations)
actor critic · policy actor · critic algorithms · 195 papers
Approaches in this cluster:
- Actor-critic theory (76 papers)
Analyze convergence, sample efficiency, and implicit bias of actor-critic algorithms. - Off-policy actor-critic with expressive policies (71 papers)
Improve off-policy actor-critic by exploiting past success, stochastic policies, and value-policy links. - Stabilizing deep actor-critic (48 papers)
Study implementation choices, overestimation, and normalization to make deep actor-critic reliable.
Most cited and most cited since 2024:
- Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor (ICML 2018 · 3,754 citations)
- Addressing Function Approximation Error in Actor-Critic Methods (ICML 2018 · 2,316 citations)
- Diffusion Actor-Critic with Entropy Regulator (NeurIPS 2024 · 10 citations)
- Two-Stage Evolutionary Reinforcement Learning for Enhancing Exploration and Exploitation (AAAI 2024 · 9 citations)
safe · constrained · constraints · 160 papers
Approaches in this cluster:
- Safety layers and shielding (85 papers)
Add safety editors, shields, or constraint-conditioned policies to keep agents within safe behavior. - Constrained policy optimization (48 papers)
Solve constrained MDPs with Lagrangian-style methods, handling bias, exploration, and model mismatch. - Safe policy improvement (27 papers)
Improve over a baseline policy with guarantees using bootstrapping, robust regret, and conformal control.
Most cited and most cited since 2024:
- Safe Model-based Reinforcement Learning with Stability Guarantees (NeurIPS 2017 · 319 citations)
- Reward Constrained Policy Optimization (ICLR 2019 · 254 citations)
- Balance Reward and Safety Optimization for Safe Reinforcement Learning: A Perspective of Gradient Manipulation (AAAI 2024 · 12 citations)
- Safe Reinforcement Learning with Instantaneous Constraints: The Role of Aggressive Exploration (AAAI 2024 · 7 citations)
ope · estimators · importance · 132 papers
Approaches in this cluster:
- Importance-weighted OPE estimators (57 papers)
Build doubly robust and marginalized importance-sampling estimators for off-policy evaluation. - Confounding-robust policy evaluation (47 papers)
Estimate policy value under confounding or distribution shift via counterfactual and density-ratio methods. - Confidence intervals and OPE benchmarks (28 papers)
Provide statistical guarantees, long-horizon methods, and risk-aware assessments for off-policy evaluation.
Most cited and most cited since 2024:
- DualDICE: Behavior-Agnostic Estimation of Discounted Stationary Distribution Corrections (NeurIPS 2019 · 111 citations)
- Breaking the Curse of Horizon: Infinite-Horizon Off-Policy Estimation (NeurIPS 2018 · 94 citations)
- Abstract Reward Processes: Leveraging State Abstraction for Consistent Off-Policy Evaluation (NeurIPS 2024 · 2 citations)
- Logarithmic Smoothing for Pessimistic Off-Policy Evaluation, Selection and Learning (NeurIPS 2024 · 2 citations)
Related topics in Reinforcement learning
- MDPs and planning (682)
- Visual and representation RL (2,133)
- Imitation learning (431)
- RL from feedback (737)
- Offline RL (623)
