Research map / Multi-agent systems and games
LLM agents: research map
1,232 accepted papers on LLM agents in Multi-agent systems and games, from ICML, NeurIPS, ICLR, CVPR and AAAI (2016–2026), grouped into 2 clusters and 9 approaches. The busiest year so far is 2026.
Within Multi-agent systems and games, its share grew from 10.3% in 2023–24 to 60.4% in 2025–26 (65 → 1,157 papers at ICML, NeurIPS, CVPR and AAAI, the venues with data for all four years).
Explore LLM agents in the interactive map
Working on something in this topic? Describe your idea in scime atlas to see which approach it falls under, the closest papers by meaning and how crowded the spot has become.
Approaches and key papers
gui · bench · software · 762 papers
Approaches in this cluster:
- LLM multi-agent decision making (194 papers)
Designs collaboration, hierarchy, and credit assignment for LLM-based multi-agent teams. - Self-evolving language agent learning (188 papers)
Improves language agents via search, step-level value models, reinforcement learning, and self-distillation. - Robust and efficient agent execution (174 papers)
Benchmarks tool-using agent robustness and speeds execution with plan caching and compilation. - Agent data synthesis and behavior studies (108 papers)
Synthesizes agent training data and examines behavioral biases in interactive LLM agents. - Agent safety and attack benchmarks (98 papers)
Benchmarks multi-turn attacks, prompt injection, and malicious tool use against LLM agents.
Most cited and most cited since 2024:
- Reflexion: language agents with verbal reinforcement learning (NeurIPS 2023 · 559 citations)
- MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework (ICLR 2024 · 153 citations)
- AgentBench: Evaluating LLMs as Agents (ICLR 2024 · 53 citations)
- AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge Bases (NeurIPS 2024 · 49 citations)
llm agents · llms · agent systems · 470 papers
Approaches in this cluster:
- Agentic benchmarks and orchestration (185 papers)
Benchmarks long-horizon agent memory, reliability, and orchestration of sub-agents. - LLM multi-agent collaboration and safety (84 papers)
Studies collaboration, competition, evaluation, and monitoring of LLM-based multi-agent systems. - Software engineering agent benchmarks (129 papers)
Builds repository-level environments and benchmarks for coding and computer-use agents. - GUI agent memory and trajectories (72 papers)
Improves GUI agents through memory, synthesized interaction data, and desktop benchmarks.
Most cited and most cited since 2024:
- Mind2Web: Towards a Generalist Agent for the Web (NeurIPS 2023 · 69 citations)
- ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate (ICLR 2024 · 55 citations)
- VerilogCoder: Autonomous Verilog Coding Agents with Graph-based Planning and Abstract Syntax Tree (AST)-based Waveform Tracing Tool (AAAI 2025 · 46 citations)
- OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments (NeurIPS 2024 · 39 citations)
Related topics in Multi-agent systems and games
- Cooperative multi-agent RL (593)
- Games and equilibria (526)
- Human-AI and goal-conditioned (1,160)
