Research map / Multi-agent systems and games
Cooperative multi-agent RL: research map
593 accepted papers on Cooperative multi-agent RL in Multi-agent systems and games, from ICML, NeurIPS, ICLR, CVPR and AAAI (2016–2026), grouped into 2 clusters and 8 approaches. The busiest year so far is 2026.
Within Multi-agent systems and games, its share shrank from 22.4% in 2023–24 to 11.9% in 2025–26 (141 → 228 papers at ICML, NeurIPS, CVPR and AAAI, the venues with data for all four years).
Explore Cooperative multi-agent RL in the interactive map
Working on something in this topic? Describe your idea in scime atlas to see which approach it falls under, the closest papers by meaning and how crowded the spot has become.
Approaches and key papers
reinforcement marl · communication · marl algorithms · 362 papers
Approaches in this cluster:
- Cooperative MARL algorithms (149 papers)
Proposes parameter-sharing, primal-dual, and benchmarked algorithms for cooperative multi-agent reinforcement learning. - MARL benchmarks and grouping (121 papers)
Introduces benchmarks and agent grouping or experience-sharing schemes for cooperative MARL. - Learned inter-agent communication (58 papers)
Learns when, what, and how to communicate among agents under bandwidth and delay limits. - Intrinsic-reward exploration for MARL (34 papers)
Drives multi-agent exploration with individual and collective intrinsic rewards and curiosity.
Most cited and most cited since 2024:
- Learning to Communicate with Deep Multi-Agent Reinforcement Learning (NeurIPS 2016 · 1,057 citations)
- The Surprising Effectiveness of PPO in Cooperative Multi-Agent Games (NeurIPS 2022 · 597 citations)
- Carbon Footprint Reduction for Sustainable Data Centers in Real-Time (AAAI 2024 · 25 citations)
- T2MAC: Targeted and Trusted Multi-Agent Communication through Selective Engagement and Evidence-Driven Integration (AAAI 2024 · 18 citations)
decentralized · value · joint · 231 papers
Approaches in this cluster:
- Decentralized actor-critic MARL (77 papers)
Trains decentralized or centralized-critic policy-gradient agents with limited communication. - Offline and networked MARL (65 papers)
Learns multi-agent policies from fixed datasets or with networked, scalable value decomposition. - Coordination graphs (50 papers)
Models agent interactions with sparse coordination graphs and search-based planning. - Value factorization (39 papers)
Decomposes joint Q-functions into per-agent values with theory and improved factorization schemes.
Most cited and most cited since 2024:
- Multi-Agent Actor-Critic for Mixed Cooperative-Competitive Environments (NeurIPS 2017 · 1,002 citations)
- Stabilising Experience Replay for Deep Multi-Agent Reinforcement Learning (ICML 2017 · 420 citations)
- Intrinsic Action Tendency Consistency for Cooperative Multi-Agent Reinforcement Learning (AAAI 2024 · 9 citations)
- STAS: Spatial-Temporal Return Decomposition for Solving Sparse Rewards Problems in Multi-agent Reinforcement Learning (AAAI 2024 · 9 citations)
Related topics in Multi-agent systems and games
- LLM agents (1,232)
- Games and equilibria (526)
- Human-AI and goal-conditioned (1,160)
