atlas

Research map / Large language models

Language, safety and interpretability: research map

2,928 accepted papers on Language, safety and interpretability in Large language models, from ICML, NeurIPS, ICLR, CVPR and AAAI (2016–2026), grouped into 6 clusters and 26 approaches. The busiest year so far is 2026.

Within Large language models, its share shrank from 35.2% in 2023–24 to 26.9% in 2025–26 (700 → 1,847 papers at ICML, NeurIPS, CVPR and AAAI, the venues with data for all four years).

2016: 10162017: 10172018: 41182019: 52192020: 75202021: 92212022: 101222023: 183232024: 517242025: 833252026: 101426

Explore Language, safety and interpretability in the interactive map

Working on something in this topic? Describe your idea in scime atlas to see which approach it falls under, the closest papers by meaning and how crowded the spot has become.

Approaches and key papers

prompt · uncertainty · shot · 1,303 papers

Approaches in this cluster:

Most cited and most cited since 2024:

languages · pretraining · multilingual · 456 papers

Approaches in this cluster:

Most cited and most cited since 2024:

natural language · speech · generation · 378 papers

Approaches in this cluster:

Most cited and most cited since 2024:

safety · jailbreak · harmful · 340 papers

Approaches in this cluster:

Most cited and most cited since 2024:

watermarking · generated text · watermarks · 228 papers

Approaches in this cluster:

Most cited and most cited since 2024:

representations · interpretability · heads · 223 papers

Approaches in this cluster:

Most cited and most cited since 2024:

Related topics in Large language models