Research map / Reinforcement learning
Visual and representation RL: research map
2,133 accepted papers on Visual and representation RL in Reinforcement learning, from ICML, NeurIPS, ICLR, CVPR and AAAI (2016–2026), grouped into 6 clusters and 23 approaches. The busiest year so far is 2026.
Within Reinforcement learning, its share shrank from 36.5% in 2023–24 to 30.0% in 2025–26 (572 → 696 papers at ICML, NeurIPS, CVPR and AAAI, the venues with data for all four years).
Explore Visual and representation RL in the interactive map
Working on something in this topic? Describe your idea in scime atlas to see which approach it falls under, the closest papers by meaning and how crowded the spot has become.
Approaches and key papers
search · decision · systems · 1,011 papers
Approaches in this cluster:
- In-context and adaptive RL (202 papers)
Enable fast adaptation in RL via in-context learning, memory, and domain transfer. - Scalable RL systems and environments (179 papers)
Scale RL training and benchmark environments. - Probabilistic and autonomous RL (177 papers)
Frame RL as inference and learn without episodic resets. - RL for combinatorial routing (158 papers)
Solve routing and NP-hard problems with learned RL policies and search. - RL with tree search for discovery (121 papers)
Search molecules and plans with RL and Monte Carlo tree search. - Inferring decision-making from behavior (93 papers)
Interpret human and animal decisions via inverse RL and policy recovery. - Curriculum and environment design (81 papers)
Generates training environments and curricula to improve agent generalization and skill acquisition.
Most cited and most cited since 2024:
- A Deep Reinforced Model for Abstractive Summarization (ICLR 2018 · 1,255 citations)
- Action-Decision Networks for Visual Tracking With Deep Reinforcement Learning (CVPR 2017 · 574 citations)
- Prompt to Transfer: Sim-to-Real Transfer for Traffic Signal Control with Prompt Learning (AAAI 2024 · 41 citations)
- EarnHFT: Efficient Hierarchical Reinforcement Learning for High Frequency Trading (AAAI 2024 · 21 citations)
representation · latent · generalization · 329 papers
Approaches in this cluster:
- Learned state and action representations (148 papers)
Learn compact representations of states, actions, or contexts to speed up and generalize RL. - Robust visual encoders for RL (86 papers)
Make pixel-based RL generalize to visual perturbations via pretrained encoders, invariance, and attention. - Latent-space world models (83 papers)
Learn latent dynamics models and temporal-consistency objectives for model-based and transfer RL. - Successor features for transfer (12 papers)
Use successor representations and features to reuse value information across tasks and policies.
Most cited and most cited since 2024:
- Benchmarking Deep Reinforcement Learning for Continuous Control (ICML 2016 · 949 citations)
- Recurrent World Models Facilitate Policy Evolution (NeurIPS 2018 · 452 citations)
- Jointly Training and Pruning CNNs via Learnable Agent Guidance and Alignment (CVPR 2024 · 20 citations)
- What Effects the Generalization in Visual Reinforcement Learning: Policy Consistency with Truncated Return Prediction (AAAI 2024 · 9 citations)
robot · manipulation · simulation · 271 papers
Approaches in this cluster:
- Skill learning and diffusion policies (86 papers)
Learn transferable skills and diffusion policies for continuous control. - Dexterous manipulation transfer (141 papers)
Transfer manipulation skills with residual learning, benchmarks, and cross-embodiment data. - Vision-based robot RL (44 papers)
Learn transferable visual control policies and world models for embodied agents.
Most cited and most cited since 2024:
- SAPIEN: A SimulAted Part-Based Interactive ENvironment (CVPR 2020 · 402 citations)
- Hindsight Experience Replay (NeurIPS 2017 · 337 citations)
- Deep Reinforcement Learning for Robotics: A Survey of Real-World Successes (AAAI 2025 · 172 citations)
- Hierarchical Diffusion Policy for Kinematics-Aware Multi-Task Robotic Manipulation (CVPR 2024 · 54 citations)
exploration · intrinsic · skills · 265 papers
Approaches in this cluster:
- Novelty and impact-driven exploration (85 papers)
Reward visiting novel or impactful states, or diversify behavior, to explore better. - Hierarchical skills and priors (84 papers)
Learn reusable sub-policies and skill priors for hierarchical reinforcement learning. - Reward shaping for sparse rewards (51 papers)
Shape or learn intrinsic rewards and curricula to cope with sparse reward signals. - Unsupervised skill discovery (45 papers)
Discover diverse skills without reward using mutual-information and contrastive objectives.
Most cited and most cited since 2024:
- Hierarchical Deep Reinforcement Learning: Integrating Temporal Abstraction and Intrinsic Motivation (NeurIPS 2016 · 822 citations)
- Parameter Space Noise for Exploration (ICLR 2018 · 359 citations)
- Robust Policy Learning via Offline Skill Diffusion (AAAI 2024 · 6 citations)
- Bootstrapped Reward Shaping (AAAI 2025 · 5 citations)
goal · conditioned · hierarchical · 156 papers
Approaches in this cluster:
- Goal-conditioned policy learning (66 papers)
Train goal-reaching policies with learned goal representations, discriminative rewards, and pretraining. - Graph search and compositional goals (55 papers)
Plan over goal graphs or compositional goal abstractions to improve goal-conditioned RL. - Subgoal generation for hierarchical RL (35 papers)
Generate adjacency-constrained or probabilistic subgoals so high-level policies guide low-level control.
Most cited and most cited since 2024:
- Visual Reinforcement Learning with Imagined Goals (NeurIPS 2018 · 307 citations)
- Data-Efficient Hierarchical Reinforcement Learning (NeurIPS 2018 · 248 citations)
- Unveiling the Significance of Toddler-Inspired Reward Transition in Goal-Oriented Reinforcement Learning (AAAI 2024 · 3 citations)
- Goal Conditioned Reinforcement Learning for Photo Finishing Tuning (NeurIPS 2024 · 3 citations)
meta · adaptation · context · 101 papers
Approaches in this cluster:
- Task-inference meta-RL (57 papers)
Meta-train context encoders and task selection so agents infer new tasks and adapt quickly. - Task structure and curricula for meta-RL (44 papers)
Exploit subtask structure, curricula, and sampling choices to improve meta-RL adaptation.
Most cited and most cited since 2024:
- Efficient Off-Policy Meta-Reinforcement Learning via Probabilistic Context Variables (ICML 2019 · 214 citations)
- Meta-Reinforcement Learning of Structured Exploration Strategies (NeurIPS 2018 · 174 citations)
- Generalizable Task Representation Learning for Offline Meta-Reinforcement Learning with Data Limitations (AAAI 2024 · 12 citations)
- Meta-Reinforcement Learning via Exploratory Task Clustering (AAAI 2024 · 7 citations)
Related topics in Reinforcement learning
- MDPs and planning (682)
- Policy gradient and actor-critic (1,502)
- Imitation learning (431)
- RL from feedback (737)
- Offline RL (623)
