Large language models: research map
Reasoning, fine-tuning, retrieval, agents, alignment and safety of LLMs.
9,634 accepted papers at ICML, NeurIPS, ICLR, CVPR and AAAI (2016–2026), in 7 topics. Within all five venues, its share grew from 7.5% in 2023–24 to 15.1% in 2025–26 (1,986 → 6,866 papers at ICML, NeurIPS, CVPR and AAAI, the venues with data for all four years).
Explore Large language models in the interactive map
Topics
- Language, safety and interpretability · 2,928 papers
LLM unlearning and benchmarks, LLM pruning and scaling data, Controlled text generation - Agents, code and math · 1,722 papers
Mathematical and structural reasoning benchmarks, Inductive and abstract reasoning evaluation, LLMs for optimization and heuristics - Reasoning and chain of thought · 1,557 papers
Efficient and adaptive thinking, Test-time compute scaling, RL with verifiable rewards for reasoning - Retrieval and knowledge · 1,064 papers
Knowledge graph reasoning with LLMs, Knowledge editing in LLMs, Knowledge-augmented question answering - Instruction and fine-tuning · 871 papers
Efficient LLM fine-tuning systems, Fine-tuning generalization and delta compression, Fine-tuning for reasoning data - AI and society · 827 papers
Human-centered trust and agency, Social-science view of AI evaluation, Data commons and human baselines - Alignment and preferences · 665 papers
Steering and correcting LLM behavior, Reward modeling for diverse preferences, LLM-as-judge and feedback data
Most cited papers in Large language models
- GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding (ICLR 2019 · 4,101 citations)
- Language Models are Few-Shot Learners (NeurIPS 2020 · 2,966 citations)
- Chain-of-Thought Prompting Elicits Reasoning in Large Language Models (NeurIPS 2022 · 2,251 citations)
- XLNet: Generalized Autoregressive Pretraining for Language Understanding (NeurIPS 2019 · 1,824 citations)
- Training language models to follow instructions with human feedback (NeurIPS 2022 · 1,219 citations)
- The Curious Case of Neural Text Degeneration (ICLR 2020 · 1,122 citations)
- SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems (NeurIPS 2019 · 984 citations)
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena (NeurIPS 2023 · 810 citations)
- Self-Consistency Improves Chain of Thought Reasoning in Language Models (ICLR 2023 · 706 citations)
- Wizard of Wikipedia: Knowledge-Powered Conversational Agents (ICLR 2019 · 623 citations)
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model (NeurIPS 2023 · 418 citations)
- CodeXGLUE: A Machine Learning Benchmark Dataset for Code Understanding and Generation (NeurIPS 2021 · 408 citations)
