Research map / Large language models
Reasoning and chain of thought: research map
1,557 accepted papers on Reasoning and chain of thought in Large language models, from ICML, NeurIPS, ICLR, CVPR and AAAI (2016–2026), grouped into 3 clusters and 12 approaches. The busiest year so far is 2026.
Within Large language models, its share grew from 6.5% in 2023–24 to 20.6% in 2025–26 (130 → 1,417 papers at ICML, NeurIPS, CVPR and AAAI, the venues with data for all four years).
Explore Reasoning and chain of thought in the interactive map
Working on something in this topic? Describe your idea in scime atlas to see which approach it falls under, the closest papers by meaning and how crowded the spot has become.
Approaches and key papers
test time · scaling · rl · 880 papers
Approaches in this cluster:
- Efficient and adaptive thinking (231 papers)
Control reasoning length and self-verification so models think wisely rather than longer. - Test-time compute scaling (205 papers)
Allocate inference-time compute through confidence, branching, and self-consistency strategies. - RL with verifiable rewards for reasoning (270 papers)
Train reasoning LLMs with reinforcement learning and analyze mathematical reasoning gains. - Routing, pruning and distillation for reasoning (135 papers)
Reduce reasoning cost via model routing, small-large collaboration, pruning, and distillation. - Process reward models (39 papers)
Train process-level reward models with synthetic annotation and calibration to verify reasoning steps.
Most cited and most cited since 2024:
- Measuring Mathematical Problem Solving With the MATH Dataset (NeurIPS 2021 · 276 citations)
- Inference-Time Intervention: Eliciting Truthful Answers from a Language Model (NeurIPS 2023 · 56 citations)
- The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity (NeurIPS 2025 · 29 citations)
- ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search (NeurIPS 2024 · 21 citations)
chain thought · thought cot · cot reasoning · 414 papers
Approaches in this cluster:
- Chain-of-thought training and analysis (127 papers)
Train and analyze chain-of-thought reasoning, including its faithfulness and information content. - Long CoT dynamics and efficiency (136 papers)
Study the geometry, legibility, and efficiency of long reasoning chains. - Search and self-correction in reasoning (123 papers)
Improve reasoning with trial-and-error, calibration, rollback, and selection-inference loops. - Theory of autoregressive reasoning (28 papers)
Analyze compositional generalization and learnability of reasoning by next-token predictors.
Most cited and most cited since 2024:
- Chain-of-Thought Prompting Elicits Reasoning in Large Language Models (NeurIPS 2022 · 2,251 citations)
- Self-Consistency Improves Chain of Thought Reasoning in Language Models (ICLR 2023 · 706 citations)
- Graph of Thoughts: Solving Elaborate Problems with Large Language Models (AAAI 2024 · 512 citations)
- MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark (NeurIPS 2024 · 82 citations)
decoding · speculative · cache · 263 papers
Approaches in this cluster:
- Efficient LLM serving (126 papers)
Schedule queries, reuse prefixes, and optimize architectures for efficient LLM inference. - Speculative decoding (92 papers)
Use draft models and acceptance-rate optimization to decode faster without changing outputs. - KV cache compression (45 papers)
Evict or merge key-value cache entries to cut memory in LLM inference.
Most cited and most cited since 2024:
- ARTrackV2: Prompting Autoregressive Tracker Where to Look and How to Describe (CVPR 2024 · 108 citations)
- H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models (NeurIPS 2023 · 66 citations)
- SnapKV: LLM Knows What You are Looking for Before Generation (NeurIPS 2024 · 51 citations)
- Hot or Cold? Adaptive Temperature Sampling for Code Generation with Large Language Models (AAAI 2024 · 34 citations)
Related topics in Large language models
- Instruction and fine-tuning (871)
- Retrieval and knowledge (1,064)
- Agents, code and math (1,722)
- Alignment and preferences (665)
- AI and society (827)
- Language, safety and interpretability (2,928)
