Research map / Large language models
Language, safety and interpretability: research map
2,928 accepted papers on Language, safety and interpretability in Large language models, from ICML, NeurIPS, ICLR, CVPR and AAAI (2016–2026), grouped into 6 clusters and 26 approaches. The busiest year so far is 2026.
Within Large language models, its share shrank from 35.2% in 2023–24 to 26.9% in 2025–26 (700 → 1,847 papers at ICML, NeurIPS, CVPR and AAAI, the venues with data for all four years).
Explore Language, safety and interpretability in the interactive map
Working on something in this topic? Describe your idea in scime atlas to see which approach it falls under, the closest papers by meaning and how crowded the spot has become.
Approaches and key papers
prompt · uncertainty · shot · 1,303 papers
Approaches in this cluster:
- LLM unlearning and benchmarks (298 papers)
Evaluate LLMs on unlearning, fallacies, and procedure generalization, with adaptation to domains. - LLM pruning and scaling data (285 papers)
Prune LLMs after training and study data effects on scaling and collapse. - Controlled text generation (220 papers)
Steer generation with plug-in attribute models, instructions, and token-level input refinement. - Faithfulness of LLM explanations (169 papers)
Evaluate whether natural-language explanations are faithful using counterfactuals and simulatability. - Data and evaluation methodology (136 papers)
Assess datasets, evaluation failures, and summarization with description-length and data-curation tools. - In-context learning analysis (95 papers)
Explain and extend in-context learning via latent-variable views, many-shot, and kNN prompting. - Parallel and early-exit decoding (100 papers)
Speed up decoding through parallel, layer-adaptive, and diffusion-based generation.
Most cited and most cited since 2024:
- GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding (ICLR 2019 · 4,101 citations)
- XLNet: Generalized Autoregressive Pretraining for Language Understanding (NeurIPS 2019 · 1,824 citations)
- Enhancing Job Recommendation through LLM-Based Generative Adversarial Networks (AAAI 2024 · 82 citations)
- Bootstrapping Large Language Models for Radiology Report Generation (AAAI 2024 · 72 citations)
languages · pretraining · multilingual · 456 papers
Approaches in this cluster:
- Pruning and circuits in LLMs (135 papers)
Prune or extract task-specific circuits and structure from large language models. - Reasoning mechanisms and failure modes (114 papers)
Analyze symbolic mechanisms, attention glitches, and hallucination in LLM reasoning. - Linguistic encoding in LLM representations (123 papers)
Probe how LLM embeddings encode syntax, semantics, and brain-aligned information. - Task vectors and steering (84 papers)
Locate task representations in attention heads and steer models via representation interventions.
Most cited and most cited since 2024:
- Incorporating Context into Language Encoding Models for fMRI (NeurIPS 2018 · 163 citations)
- Measuring abstract reasoning in neural networks (ICML 2018 · 141 citations)
- LLMs are Good Sign Language Translators (CVPR 2024 · 70 citations)
- Fluctuation-Based Adaptive Structured Pruning for Large Language Models (AAAI 2024 · 43 citations)
natural language · speech · generation · 378 papers
Approaches in this cluster:
- Text generation and evaluation metrics (79 papers)
Study embeddings, metrics, and language-model-based evaluation for natural language generation. - Speech and audio language models (111 papers)
Combine large language models with speech and audio understanding and generation. - Multilingual LLM data and culture (69 papers)
Allocate multilingual data, select pretraining data, and adapt models to cultural differences. - Sign language datasets and translation (23 papers)
Collect sign language datasets and improve translation with multilingual and monolingual data. - Sign and cross-lingual translation (96 papers)
Translates sign and spoken languages using large language models and multilingual pretraining.
Most cited and most cited since 2024:
- Improving Sign Language Translation With Monolingual Data by Sign Back-Translation (CVPR 2021 · 277 citations)
- Data Noising as Smoothing in Neural Network Language Models (ICLR 2017 · 213 citations)
- AudioGPT: Understanding and Generating Speech, Music, Sound, and Talking Head (AAAI 2024 · 114 citations)
- CultureLLM: Incorporating Cultural Differences into Large Language Models (NeurIPS 2024 · 26 citations)
safety · jailbreak · harmful · 340 papers
Approaches in this cluster:
- Uncertainty quantification for LLMs (136 papers)
Quantify confidence and uncertainty to benchmark and evaluate language models. - Statistical frameworks for LLM evaluation (124 papers)
Use conformal prediction, ranking models, and noise analysis for valid LLM evaluation. - Bias measurement and mitigation (59 papers)
Measure and reduce gender, geographic, and implicit bias in LLMs. - Jailbreak elicitation and risk (21 papers)
Elicit harmful outputs through interaction, many-shot, and weak-to-strong jailbreaks.
Most cited and most cited since 2024:
- The Curious Case of Neural Text Degeneration (ICLR 2020 · 1,122 citations)
- Towards Understanding and Mitigating Social Biases in Language Models (ICML 2021 · 124 citations)
- Improving Factuality and Reasoning in Language Models through Multiagent Debate (ICML 2024 · 94 citations)
- Membership Inference Attacks against Fine-tuned Large Language Models via Self-prompt Calibration (NeurIPS 2024 · 28 citations)
watermarking · generated text · watermarks · 228 papers
Approaches in this cluster:
- Text watermarking (80 papers)
Embed detectable, robust watermarks into LLM output with provable guarantees. - Zero-shot machine-text detection (82 papers)
Detect LLM-generated text with perplexity-based and reward-based zero-shot scores. - Data and hallucination detection (66 papers)
Detect pretraining data, hallucinations, and unsafe training data using token-level and attribution signals.
Most cited and most cited since 2024:
- Bad Actor, Good Advisor: Exploring the Role of Large Language Models in Fake News Detection (AAAI 2024 · 221 citations)
- A Holistic Approach to Undesired Content Detection in the Real World (AAAI 2023 · 119 citations)
- Position: Will we run out of data? Limits of LLM scaling based on human-generated data (ICML 2024 · 83 citations)
- Position: On the Possibilities of AI-Generated Text Detection (ICML 2024 · 53 citations)
representations · interpretability · heads · 223 papers
Approaches in this cluster:
- Safety alignment of LLMs (123 papers)
Study how safety training, safeguards, and unsafe-data forgetting shape LLM behavior. - Mechanistic analysis of attention and representations (43 papers)
Trace attention heads and representation editing to explain in-context learning and model behavior. - Jailbreak attacks and defenses (57 papers)
Craft automated, multi-turn, and out-of-distribution jailbreaks and defend against them.
Most cited and most cited since 2024:
- Jailbroken: How Does LLM Safety Training Fail? (NeurIPS 2023 · 130 citations)
- Anti-efficient encoding in emergent communication (NeurIPS 2019 · 68 citations)
- Tree of Attacks: Jailbreaking Black-Box LLMs Automatically (NeurIPS 2024 · 55 citations)
- AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models (ICLR 2024 · 27 citations)
Related topics in Large language models
- Reasoning and chain of thought (1,557)
- Instruction and fine-tuning (871)
- Retrieval and knowledge (1,064)
- Agents, code and math (1,722)
- Alignment and preferences (665)
- AI and society (827)
