Research map / Architectures and efficiency
Mixture of experts: research map
300 accepted papers on Mixture of experts in Architectures and efficiency, from ICML, NeurIPS, ICLR, CVPR and AAAI (2016–2026), grouped into 2 clusters and 5 approaches. The busiest year so far is 2026.
Within Architectures and efficiency, its share grew from 2.0% in 2023–24 to 6.7% in 2025–26 (39 → 239 papers at ICML, NeurIPS, CVPR and AAAI, the venues with data for all four years).
Explore Mixture of experts in the interactive map
Working on something in this topic? Describe your idea in scime atlas to see which approach it falls under, the closest papers by meaning and how crowded the spot has become.
Approaches and key papers
sparse · scaling · moes · 267 papers
Approaches in this cluster:
- Expert routing and specialization (150 papers)
Improves expert specialization, soft routing, and interpretation in mixture-of-experts models. - Sparse MoE as regularization and attention (45 papers)
Applies sparse expert layers to attention, dropout-style regularization, and sparsely activated transformers. - Efficient MoE language model systems (72 papers)
Restructures LLMs into MoE and speeds distributed MoE training with routing and scheduling.
Most cited and most cited since 2024:
- GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding (ICLR 2021 · 351 citations)
- Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer (ICLR 2017 · 261 citations)
- Complexity Experts are Task-Discriminative Learners for Any Image Restoration (CVPR 2025 · 51 citations)
- From Sparse to Soft Mixtures of Experts (ICLR 2024 · 25 citations)
lora · rank · tuning · 33 papers
Approaches in this cluster:
- Mixture-of-experts LoRA (18 papers)
Combines low-rank adapters with expert routing for multi-task and parameter-efficient adaptation. - MoE LLM compression (15 papers)
Compresses mixture-of-experts LLMs through decomposition, delta decompression, and expert-level quantization.
Most cited and most cited since 2024:
- Expert Gate: Lifelong Learning With a Network of Experts (CVPR 2017 · 613 citations)
- Multi-Task Dense Prediction via Mixture of Low-Rank Experts (CVPR 2024 · 50 citations)
- SPMTrack: Spatio-Temporal Parameter-Efficient Fine-Tuning with Mixture of Experts for Scalable Visual Tracking (CVPR 2025 · 16 citations)
- Pushing Mixture of Experts to the Limit: Extremely Parameter Efficient MoE for Instruction Tuning (ICLR 2024 · 11 citations)
Related topics in Architectures and efficiency
- Pruning and efficient tuning (1,147)
- Attention and transformers (2,296)
- Neural architecture search (380)
- Quantization (646)
- RNNs and normalization (2,827)
- Spiking networks (289)
- CNN design (853)
