Research map / Architectures and efficiency
Attention and transformers: research map
2,296 accepted papers on Attention and transformers in Architectures and efficiency, from ICML, NeurIPS, ICLR, CVPR and AAAI (2016–2026), grouped into 6 clusters and 25 approaches. The busiest year so far is 2026.
Within Architectures and efficiency, its share grew from 29.8% in 2023–24 to 32.9% in 2025–26 (583 → 1,166 papers at ICML, NeurIPS, CVPR and AAAI, the venues with data for all four years).
Explore Attention and transformers in the interactive map
Working on something in this topic? Describe your idea in scime atlas to see which approach it falls under, the closest papers by meaning and how crowded the spot has become.
Approaches and key papers
kv · cache · long context · 531 papers
Approaches in this cluster:
- Sparse and structured attention (163 papers)
Redesigns self-attention with sparsity, structure, and rank analysis for efficiency and interpretability. - Attention head and block redesign (127 papers)
Modifies transformer attention heads, mixing, and primal-dual views to improve learning and efficiency. - Efficient image transformers (84 papers)
Applies sparse or shared attention and U-shaped transformers to image restoration and matching. - Long-context KV and memory attention (88 papers)
Speeds long-context LLM inference with hierarchical sparse attention and key-value memory designs. - Transformer neural operators for PDEs (69 papers)
Uses attention-based neural operators to simulate and pre-train on partial differential equations.
Most cited and most cited since 2024:
- Uformer: A General U-Shaped Transformer for Image Restoration (CVPR 2022 · 2,236 citations)
- A STRUCTURED SELF-ATTENTIVE SENTENCE EMBEDDING (ICLR 2017 · 1,630 citations)
- Adapt or Perish: Adaptive Sparse Transformer with Attentive Feature Refinement for Image Restoration (CVPR 2024 · 210 citations)
- Efficient LoFTR: Semi-Dense Local Feature Matching with Sparse-Like Speed (CVPR 2024 · 177 citations)
vision · vit · transformers · 384 papers
Approaches in this cluster:
- Transformer expressivity and learnability (217 papers)
Theoretically analyzes what algorithms and structured data transformers can represent and learn. - Attention head mechanisms (75 papers)
Studies attention heads, induction circuits, and positional encodings behind length generalization. - Training dynamics of in-context learning (62 papers)
Proves how transformers acquire in-context learning and induction heads during training. - Vision transformer design (30 papers)
Improves vision transformer robustness, efficiency, and training through architectural changes.
Most cited and most cited since 2024:
- Conditional Positional Encodings for Vision Transformers (ICLR 2023 · 406 citations)
- Intriguing Properties of Vision Transformers (NeurIPS 2021 · 302 citations)
- You Only Need Less Attention at Each Stage in Vision Transformers (CVPR 2024 · 21 citations)
- Mamba-Reg: Vision Mamba Also Needs Registers (CVPR 2025 · 19 citations)
self attention · attention mechanism · softmax · 369 papers
Approaches in this cluster:
- Efficient and robust vision transformers (124 papers)
Improves vision transformer efficiency, robustness, and registers through adaptive computation and architecture changes. - Hybrid and interpretable ViT designs (94 papers)
Enhances vision transformers with convolutional features, local attention, and interpretable decoders. - Token merging and masked pretraining (82 papers)
Trains vision transformers with token labeling, masked image modeling, and distilled tokens. - Theory of self-attention dynamics (43 papers)
Analyzes how self-attention learns, clusters tokens, and relates to causal structure. - Vision transformer pruning (26 papers)
Prunes ViT depth, tokens, and channels using saliency and structure-aware optimization.
Most cited and most cited since 2024:
- Masked Autoencoders Are Scalable Vision Learners (CVPR 2022 · 7,924 citations)
- MetaFormer Is Actually What You Need for Vision (CVPR 2022 · 1,239 citations)
- RMT: Retentive Networks Meet Vision Transformers (CVPR 2024 · 229 citations)
- SHViT: Single-Head Vision Transformer with Memory Efficient Macro Design (CVPR 2024 · 179 citations)
sequence · long · modeling · 364 papers
Approaches in this cluster:
- State space and recurrent sequence models (172 papers)
Builds SSM, long-convolution, and recurrent models for efficient long-sequence modeling. - Sparse efficient transformers for long context (76 papers)
Reduces attention cost via sparse, hierarchical, or recurrent-memory designs for long sequences. - Linear attention approximations (76 papers)
Approximates softmax attention with kernel, linear, or gated attention for linear-time complexity. - Mamba architectures (40 papers)
Analyzes and adapts Mamba selective state space models, especially for vision tasks.
Most cited and most cited since 2024:
- Global Context-Aware Attention LSTM Networks for 3D Action Recognition (CVPR 2017 · 705 citations)
- Ask Me Anything: Dynamic Memory Networks for Natural Language Processing (ICML 2016 · 628 citations)
- MambaOut: Do We Really Need Mamba for Vision? (CVPR 2025 · 188 citations)
- FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning (ICLR 2024 · 143 citations)
transformers · language · icl · 349 papers
Approaches in this cluster:
- Transformer block and depth design (89 papers)
Reworks transformer layers with depth-weighted connections, adaptive depth, and dynamical analyses. - Efficient language transformers (105 papers)
Searches and streamlines transformer architectures for language modeling, translation, and stable training. - Language model pretraining and induction (99 papers)
Studies positional encodings, embeddings, and induction heads in language-model pretraining. - Transformer expressivity and circuits (56 papers)
Analyzes Turing completeness, induction-head representability, and attention-head stability in transformers.
Most cited and most cited since 2024:
- Meshed-Memory Transformer for Image Captioning (CVPR 2020 · 1,212 citations)
- Language Modeling with Gated Convolutional Networks (ICML 2017 · 1,114 citations)
- SaProt: Protein Language Modeling with Structure-aware Vocabulary (ICLR 2024 · 252 citations)
- DNABERT-2: Efficient Foundation Model and Benchmark For Multi-Species Genomes (ICLR 2024 · 173 citations)
global · local · spatial · 299 papers
Approaches in this cluster:
- Sparse attention for long-context LLMs (147 papers)
Selects or compresses tokens and keys to cut attention cost in long-context LLM inference. - Block-sparse attention kernels (60 papers)
Accelerates attention with block-sparse patterns and hardware-efficient kernels for LLMs and video generation. - KV cache eviction and compression (92 papers)
Prunes and compresses key-value caches with adaptive budgets to cut LLM inference memory.
Most cited and most cited since 2024:
- CSWin Transformer: A General Vision Transformer Backbone With Cross-Shaped Windows (CVPR 2022 · 1,316 citations)
- MAT: Mask-Aware Transformer for Large Hole Image Inpainting (CVPR 2022 · 440 citations)
- MobileMamba: Lightweight Multi-Receptive Visual Mamba Network (CVPR 2025 · 91 citations)
- FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision (NeurIPS 2024 · 71 citations)
Related topics in Architectures and efficiency
- Pruning and efficient tuning (1,147)
- Neural architecture search (380)
- Quantization (646)
- RNNs and normalization (2,827)
- Spiking networks (289)
- Mixture of experts (300)
- CNN design (853)
