Research map / Large language models
Alignment and preferences: research map
665 accepted papers on Alignment and preferences in Large language models, from ICML, NeurIPS, ICLR, CVPR and AAAI (2016–2026), grouped into 2 clusters and 7 approaches. The busiest year so far is 2026.
Within Large language models, its share grew from 6.6% in 2023–24 to 7.6% in 2025–26 (131 → 522 papers at ICML, NeurIPS, CVPR and AAAI, the venues with data for all four years).
Explore Alignment and preferences in the interactive map
Working on something in this topic? Describe your idea in scime atlas to see which approach it falls under, the closest papers by meaning and how crowded the spot has become.
Approaches and key papers
feedback · human · rlhf · 416 papers
Approaches in this cluster:
- Steering and correcting LLM behavior (127 papers)
Align LLMs through representation editing, influence functions, and intrinsic corpus signals. - Reward modeling for diverse preferences (90 papers)
Improve RLHF reward models with language feedback, diverse preferences, and error analysis. - LLM-as-judge and feedback data (111 papers)
Collect scaled AI feedback and study human-versus-LLM judgment in evaluation. - Reward model design and preference models (88 papers)
Combine, actively label, and generalize reward and preference models beyond Bradley-Terry.
Most cited and most cited since 2024:
- Training language models to follow instructions with human feedback (NeurIPS 2022 · 1,219 citations)
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena (NeurIPS 2023 · 810 citations)
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference (ICML 2024 · 103 citations)
- RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback (ICML 2024 · 69 citations)
preference optimization · dpo · direct preference · 249 papers
Approaches in this cluster:
- Robust and calibrated DPO (118 papers)
Make direct preference optimization robust to noisy preferences and distribution shift through calibration and data selection. - Active and token-level preference learning (58 papers)
Extend DPO with token-level objectives, reward augmentation, and active preference collection. - DPO design choices (73 papers)
Study data, scheduling, dynamic beta, and multi-reference variants of direct preference optimization.
Most cited and most cited since 2024:
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model (NeurIPS 2023 · 418 citations)
- Diffusion Model Alignment Using Direct Preference Optimization (CVPR 2024 · 154 citations)
- SimPO: Simple Preference Optimization with a Reference-Free Reward (NeurIPS 2024 · 39 citations)
- On Softmax Direct Preference Optimization for Recommendation (NeurIPS 2024 · 23 citations)
Related topics in Large language models
- Reasoning and chain of thought (1,557)
- Instruction and fine-tuning (871)
- Retrieval and knowledge (1,064)
- Agents, code and math (1,722)
- AI and society (827)
- Language, safety and interpretability (2,928)
