Research map / Label-efficient and robust learning
Knowledge distillation: research map
544 accepted papers on Knowledge distillation in Label-efficient and robust learning, from ICML, NeurIPS, ICLR, CVPR and AAAI (2016–2026), grouped into 2 clusters and 8 approaches. The busiest year so far is 2026.
Within Label-efficient and robust learning, its share grew from 4.9% in 2023–24 to 6.9% in 2025–26 (142 → 212 papers at ICML, NeurIPS, CVPR and AAAI, the venues with data for all four years).
Explore Knowledge distillation in the interactive map
Working on something in this topic? Describe your idea in scime atlas to see which approach it falls under, the closest papers by meaning and how crowded the spot has become.
Approaches and key papers
teacher student · kd · teachers · 401 papers
Approaches in this cluster:
- Rethinking distillation objectives (160 papers)
Revisits distillation losses via label smoothing, probability mass allocation and empirical analysis. - Distillation with causal and adaptive guidance (82 papers)
Improves distillation using causal intervention, dual guidance and out-of-domain data. - Few-shot and metric distillation (80 papers)
Distills compact students with limited data using grafting, regression and metric learning. - Teacher properties and robustness (54 papers)
Studies what makes teachers effective, including adversarial, random and noisy-input teachers. - Knowledge distillation for detectors (25 papers)
Distill detectors by mimicking features, rankings, and focal or global knowledge.
Most cited and most cited since 2024:
- Self-Training With Noisy Student Improves ImageNet Classification (CVPR 2020 · 2,328 citations)
- Deep Mutual Learning (CVPR 2018 · 1,819 citations)
- Logit Standardization in Knowledge Distillation (CVPR 2024 · 243 citations)
- CLIP-KD: An Empirical Study of CLIP Model Distillation (CVPR 2024 · 65 citations)
distilled · ensemble · synthetic · 143 papers
Approaches in this cluster:
- Patient consistent distillation (68 papers)
Distills with consistent teachers, relational knowledge and soft labels, including dataset distillation. - Self and data-free distillation (57 papers)
Uses self-distillation, data-free settings, decoupled losses and ensemble distillation for efficient students. - Distillation for machine translation (18 papers)
Applies selective and uncertainty-aware distillation to non-autoregressive and simultaneous translation.
Most cited and most cited since 2024:
- Decoupled Knowledge Distillation (CVPR 2022 · 960 citations)
- When does label smoothing help? (NeurIPS 2019 · 879 citations)
- Decoupled Kullback-Leibler Divergence Loss (NeurIPS 2024 · 41 citations)
- Understanding the Role of the Projector in Knowledge Distillation (AAAI 2024 · 37 citations)
Related topics in Label-efficient and robust learning
- Noisy labels (462)
- Zero-/few-shot and meta-learning (784)
- Test-time adaptation (359)
- OOD and anomaly detection (614)
- Recognition and long-tail (4,953)
- Semi-supervised learning (623)
- Self-supervised and contrastive (988)
- Domain adaptation and generalization (1,330)
