Research map / Reinforcement learning
Offline RL: research map
623 accepted papers on Offline RL in Reinforcement learning, from ICML, NeurIPS, ICLR, CVPR and AAAI (2016–2026), grouped into 2 clusters and 8 approaches. The busiest year so far is 2026.
Within Reinforcement learning, its share shrank from 15.3% in 2023–24 to 10.7% in 2025–26 (240 → 247 papers at ICML, NeurIPS, CVPR and AAAI, the venues with data for all four years).
Explore Offline RL in the interactive map
Working on something in this topic? Describe your idea in scime atlas to see which approach it falls under, the closest papers by meaning and how crowded the spot has become.
Approaches and key papers
offline reinforcement · offline rl · conservative · 394 papers
Approaches in this cluster:
- Offline RL sample complexity theory (156 papers)
Prove sample-complexity bounds for offline RL under limited data coverage. - Simple practical offline RL (144 papers)
Improve offline RL with minimalist algorithms, reweighting, and symmetry-based data efficiency. - Pessimism and uncertainty in offline RL (59 papers)
Penalize uncertain values with ensembles, Bayesian models, and pessimistic value iteration. - Behavior regularization for offline RL (35 papers)
Constrain learned policies toward the dataset with divergence penalties and in-sample learning.
Most cited and most cited since 2024:
- Conservative Q-Learning for Offline Reinforcement Learning (NeurIPS 2020 · 556 citations)
- MOPO: Model-based Offline Policy Optimization (NeurIPS 2020 · 238 citations)
- Beyond OOD State Actions: Supported Cross-Domain Offline Reinforcement Learning (AAAI 2024 · 7 citations)
- Entropy-regularized Diffusion Policy with Q-Ensembles for Offline Reinforcement Learning (NeurIPS 2024 · 7 citations)
offline online · transformer · online reinforcement · 229 papers
Approaches in this cluster:
- Offline-to-online fine-tuning (99 papers)
Pretrain on offline data then fine-tune online with calibration and balanced exploration. - Diffusion models for offline decision making (82 papers)
Use diffusion sampling and representation learning for offline and offline-to-online decisions. - Decision Transformer sequence models (40 papers)
Cast offline RL as sequence modeling with transformers, enhanced by value guidance. - Offline policy evaluation and selection (8 papers)
Evaluate and select policies from offline data using augmentation and re-weighting.
Most cited and most cited since 2024:
- Decision Transformer: Reinforcement Learning via Sequence Modeling (NeurIPS 2021 · 464 citations)
- A General Offline Reinforcement Learning Framework for Interactive Recommendation (AAAI 2021 · 63 citations)
- A Perspective of Q-value Estimation on Offline-to-Online Reinforcement Learning (AAAI 2024 · 19 citations)
- Critic-Guided Decision Transformer for Offline Reinforcement Learning (AAAI 2024 · 19 citations)
Related topics in Reinforcement learning
- MDPs and planning (682)
- Visual and representation RL (2,133)
- Policy gradient and actor-critic (1,502)
- Imitation learning (431)
- RL from feedback (737)
