Research map / Reinforcement learning
RL from feedback: research map
737 accepted papers on RL from feedback in Reinforcement learning, from ICML, NeurIPS, ICLR, CVPR and AAAI (2016–2026), grouped into 2 clusters and 7 approaches. The busiest year so far is 2026.
Within Reinforcement learning, its share grew from 5.8% in 2023–24 to 26.8% in 2025–26 (91 → 620 papers at ICML, NeurIPS, CVPR and AAAI, the venues with data for all four years).
Explore RL from feedback in the interactive map
Working on something in this topic? Describe your idea in scime atlas to see which approach it falls under, the closest papers by meaning and how crowded the spot has become.
Approaches and key papers
reasoning · grpo · llm · 506 papers
Approaches in this cluster:
- RL for LLM reasoning (187 papers)
Apply reinforcement learning to improve multi-step reasoning in language models. - Scaling RL with verifiable rewards (119 papers)
Study scaling, exploration, and learning dynamics of RL with verifiable rewards. - GRPO stability and variants (106 papers)
Analyze and fix group-relative policy optimization issues such as noise, collapse, and trust regions. - RL fine-tuning of generative models (94 papers)
Fine-tune diffusion and multimodal generative models with policy-gradient and reward objectives.
Most cited and most cited since 2024:
- Sequence Tutor: Conservative Fine-Tuning of Sequence Generation Models with KL-control (ICML 2017 · 73 citations)
- Learning to Generalize from Sparse and Underspecified Rewards (ICML 2019 · 69 citations)
- Flow-GRPO: Training Flow Matching Models via Online RL (NeurIPS 2025 · 36 citations)
- Auto MC-Reward: Automated Dense Reward Design with Large Language Models for Minecraft (CVPR 2024 · 25 citations)
human · rlhf · preferences · 231 papers
Approaches in this cluster:
- Theory of RLHF (110 papers)
Analyze sample efficiency and reward-model behavior in reinforcement learning from human feedback. - Reward modeling versus direct optimization (95 papers)
Compare reward-model RLHF with DPO-style direct policy optimization and refine preference objectives. - Preference-based RL reward learning (26 papers)
Learn rewards from pairwise preferences with improved queries, exploration, and reward-agnostic methods.
Most cited and most cited since 2024:
- Deep Reinforcement Learning from Human Preferences (NeurIPS 2017 · 500 citations)
- Interactive Learning from Policy-Dependent Human Feedback (ICML 2017 · 104 citations)
- Using Human Feedback to Fine-tune Diffusion Models without Any Reward Model (CVPR 2024 · 49 citations)
- Towards Automated RISC-V Microarchitecture Design with Reinforcement Learning (AAAI 2024 · 13 citations)
Related topics in Reinforcement learning
- MDPs and planning (682)
- Visual and representation RL (2,133)
- Policy gradient and actor-critic (1,502)
- Imitation learning (431)
- Offline RL (623)
