Research map / Large language models
Agents, code and math: research map
1,722 accepted papers on Agents, code and math in Large language models, from ICML, NeurIPS, ICLR, CVPR and AAAI (2016–2026), grouped into 4 clusters and 16 approaches. The busiest year so far is 2026.
Within Large language models, its share held steady from 18.3% in 2023–24 to 18.3% in 2025–26 (363 → 1,258 papers at ICML, NeurIPS, CVPR and AAAI, the venues with data for all four years).
Explore Agents, code and math in the interactive map
Working on something in this topic? Describe your idea in scime atlas to see which approach it falls under, the closest papers by meaning and how crowded the spot has become.
Approaches and key papers
programs · program synthesis · programming · 805 papers
Approaches in this cluster:
- Mathematical and structural reasoning benchmarks (304 papers)
Benchmark LLM math and structural reasoning, with decomposition-based analysis. - Inductive and abstract reasoning evaluation (158 papers)
Test inductive, abstract, and symbolic reasoning, including synthesized benchmarks. - LLMs for optimization and heuristics (147 papers)
Use LLMs with evolutionary search to design heuristics and model optimization problems. - Code generation benchmarks (96 papers)
Evaluate LLM code generation with evolving, domain-specific, and verification-focused benchmarks. - LLMs for scientific discovery (100 papers)
Apply LLMs to drug discovery, equation discovery, and molecular reasoning with new benchmarks.
Most cited and most cited since 2024:
- Self-Refine: Iterative Refinement with Self-Feedback (NeurIPS 2023 · 350 citations)
- Measuring Massive Multitask Language Understanding (ICLR 2021 · 325 citations)
- Can LLM Replace Stack Overflow? A Study on Robustness and Reliability of Large Language Model Code Generation (AAAI 2024 · 63 citations)
- ReEvo: Large Language Models as Hyper-Heuristics with Reflective Evolution (NeurIPS 2024 · 53 citations)
evaluating · capabilities · bench · 572 papers
Approaches in this cluster:
- Agent instruction-following benchmarks (176 papers)
Evaluate and improve LLM agents' instruction following and decision-making. - Long-horizon LLM agents (155 papers)
Train and structure LLM agents for long-horizon planning, exploration, and in-context learning. - Software engineering agents (140 papers)
Build and evaluate coding agents using execution feedback and code-graph models. - Reasoning benchmarks and platforms (47 papers)
Probe LLM reasoning with games, multilingual math, and competitive programming evaluations. - LLM-driven embodied and driving agents (54 papers)
Use language models for autonomous driving, navigation, and manipulation instructions.
Most cited and most cited since 2024:
- LMDrive: Closed-Loop End-to-End Driving with Large Language Models (CVPR 2024 · 167 citations)
- ExpeL: LLM Agents Are Experiential Learners (AAAI 2024 · 123 citations)
- ChatScene: Knowledge-Enabled Safety-Critical Scenario Generation for Autonomous Vehicles (CVPR 2024 · 79 citations)
- MDAgents: An Adaptive Collaboration of LLMs for Medical Decision-Making (NeurIPS 2024 · 58 citations)
search · optimization · decision · 258 papers
Approaches in this cluster:
- LLM-guided program synthesis (93 papers)
Synthesize programs with LLMs through decomposition, learned abstractions, and per-instance search. - Execution-based code verification (90 papers)
Verify and debug LLM-generated code using execution feedback, critics, and rigorous tests. - LLM planning with symbolic tools (44 papers)
Combine language models with action languages and commonsense knowledge for task planning. - Neural program synthesis and transpilation (31 papers)
Guide program synthesis, translation, and robot programs with learned models and probabilistic programs.
Most cited and most cited since 2024:
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation (NeurIPS 2023 · 172 citations)
- Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied Agents (ICML 2022 · 155 citations)
- SGLang: Efficient Execution of Structured Language Model Programs (NeurIPS 2024 · 109 citations)
- Large Language Models as Optimizers (ICLR 2024 · 93 citations)
theorem · proof · proving · 87 papers
Approaches in this cluster:
- Retrieval-augmented formal theorem proving (62 papers)
Prove theorems with language models using retrieval, hammers, and subgoal demonstrations. - Formal verification and proof data (25 papers)
Synthesize proof data and use autoformalization and Lean to verify LLM mathematical reasoning.
Most cited and most cited since 2024:
- HOList: An Environment for Machine Learning of Higher Order Logic Theorem Proving (ICML 2019 · 65 citations)
- GamePad: A Learning Environment for Theorem Proving (ICLR 2019 · 56 citations)
- Llemma: An Open Language Model for Mathematics (ICLR 2024 · 15 citations)
- E-GPS: Explainable Geometry Problem Solving via Top-Down Solver and Bottom-Up Generator (CVPR 2024 · 10 citations)
Related topics in Large language models
- Reasoning and chain of thought (1,557)
- Instruction and fine-tuning (871)
- Retrieval and knowledge (1,064)
- Alignment and preferences (665)
- AI and society (827)
- Language, safety and interpretability (2,928)
