Research map / Architectures and efficiency
RNNs and normalization: research map
2,827 accepted papers on RNNs and normalization in Architectures and efficiency, from ICML, NeurIPS, ICLR, CVPR and AAAI (2016–2026), grouped into 6 clusters and 28 approaches. The busiest year so far is 2026.
Within Architectures and efficiency, its share shrank from 33.3% in 2023–24 to 26.4% in 2025–26 (651 → 934 papers at ICML, NeurIPS, CVPR and AAAI, the venues with data for all four years).
Explore RNNs and normalization in the interactive map
Working on something in this topic? Describe your idea in scime atlas to see which approach it falls under, the closest papers by meaning and how crowded the spot has become.
Approaches and key papers
memory · scaling · cost · 935 papers
Approaches in this cluster:
- Efficient network compression and training (229 papers)
Accelerates and compresses networks via memory-efficient backprop, structured shrinking, and parameter prediction. - Dataset distillation and data-efficient training (158 papers)
Condenses datasets and reuses pretrained models to cut training cost. - Neural scaling laws (155 papers)
Fits scaling laws for model size, data, weight decay, and merging in language and vision. - Distributed LLM inference systems (147 papers)
Optimizes LLM serving and training with parallelism, KV reuse, batching, and GPU scheduling. - Dataset condensation and flow estimation (126 papers)
Condenses large image datasets and improves optical flow and ResNet training strategies. - Scalable approximation and equivariant compute (120 papers)
Scales Laplace approximations, equivariant attention, and tensor-program performance models.
Most cited and most cited since 2024:
- PyTorch: An Imperative Style, High-Performance Deep Learning Library (NeurIPS 2019 · 15,937 citations)
- Run, Don't Walk: Chasing Higher FLOPS for Faster Neural Networks (CVPR 2023 · 2,281 citations)
- Selective-Stereo: Adaptive Frequency Information Selection for Stereo Matching (CVPR 2024 · 103 citations)
- EquiformerV2: Improved Equivariant Transformer for Scaling to Higher-Degree Representations (ICLR 2024 · 84 citations)
layer · activation · functions · 876 papers
Approaches in this cluster:
- Theory of depth and width (223 papers)
Analyzes exploding gradients, deep linear solutions, and expressive power of deep and wide networks. - Representation geometry and collapse (169 papers)
Studies feature redundancy, neural collapse, and initialization effects on generalization in deep networks. - Layer-wise analysis of language models (145 papers)
Probes hidden representations and layer dynamics in pretrained language models. - Symmetry and equivariance in loss landscapes (95 papers)
Studies parameter symmetries and learns equivariance automatically in neural network architectures. - Mechanistic circuit interpretability (93 papers)
Identifies circuits and modular structure in networks through sparse weights and benchmarks. - Sparse dictionary representations (98 papers)
Extracts interpretable features with sparse autoencoders and dictionary views of implicit neural representations. - Width and depth of ReLU networks (53 papers)
Determines minimum width and depth needed for ReLU networks to approximate functions.
Most cited and most cited since 2024:
- Social LSTM: Human Trajectory Prediction in Crowded Spaces (CVPR 2016 · 3,627 citations)
- Learning Important Features Through Propagating Activation Differences (ICML 2017 · 2,386 citations)
- Rewrite the Stars (CVPR 2024 · 612 citations)
- Feature Reuse and Scaling: Understanding Transfer Learning with Protein Language Models (ICML 2024 · 55 citations)
end · differentiable · translation · 449 papers
Approaches in this cluster:
- Differentiable optimization layers (125 papers)
Embeds optimization problems and efficient modules as differentiable end-to-end layers. - Neural-symbolic reasoning (98 papers)
Combines recurrent networks and differentiable optimization with symbolic reasoning and program synthesis. - Non-autoregressive translation (102 papers)
Speeds machine translation decoding with non-autoregressive models and efficient decoder designs. - Learned embeddings for retrieval (76 papers)
Learns sparse, quantizable, and adaptive embeddings for efficient approximate nearest-neighbor search. - Associative memory networks (48 papers)
Builds Hopfield-style and memory-augmented networks for associative recall and external memory.
Most cited and most cited since 2024:
- Convolutional Sequence to Sequence Learning (ICML 2017 · 2,604 citations)
- Deep Speech 2 : End-to-End Speech Recognition in English and Mandarin (ICML 2016 · 2,239 citations)
- PARA-Drive: Parallelized Architecture for Real-time Autonomous Driving (CVPR 2024 · 61 citations)
- GLOP: Learning Global Partition and Local Construction for Solving Large-Scale Routing Problems in Real-Time (AAAI 2024 · 38 citations)
rnn · long · term · 303 papers
Approaches in this cluster:
- RNN training dynamics and gradients (102 papers)
Analyzes learning, generalization, and vanishing gradients in recurrent networks and proposes stable variants. - Brain-inspired recurrent architectures (87 papers)
Studies recurrent network structure, depth analogies, and biologically constrained models such as Dale's law. - Efficient recurrent architectures (60 papers)
Makes RNNs efficient via activity sparsity, dropout, variable computation, and non-normal dynamics. - LSTM design and memory (54 papers)
Improves LSTMs with tensorization, binary gates, sparse structure, and analysis of long memory.
Most cited and most cited since 2024:
- On Human Motion Prediction Using Recurrent Neural Networks (CVPR 2017 · 1,024 citations)
- Full Resolution Image Compression With Recurrent Neural Networks (CVPR 2017 · 942 citations)
- Recurrent neural networks: vanishing and exploding gradients are not the end of the story (NeurIPS 2024 · 18 citations)
- Traveling Waves Encode The Recent Past and Enhance Sequence Learning (ICLR 2024 · 10 citations)
brain · biological · feedback · 153 papers
Approaches in this cluster:
- Brain-inspired network comparison (57 papers)
Compares artificial networks with brain systems and derives biologically plausible architectures and objectives. - Hebbian and predictive coding (52 papers)
Develops Hebbian and predictive-coding learning algorithms and analyzes their dynamics. - Backprop-free credit assignment (44 papers)
Scales biologically plausible credit assignment that avoids weight symmetry and exact backpropagation.
Most cited and most cited since 2024:
- Direct Feedback Alignment Provides Learning in Deep Neural Networks (NeurIPS 2016 · 294 citations)
- Emergence of grid-like representations by training recurrent neural networks to perform spatial localization (ICLR 2018 · 203 citations)
- Joint Learning Neuronal Skeleton and Brain Circuit Topology with Permutation Invariant Encoders for Neuron Classification (AAAI 2024 · 12 citations)
- A Stable, Fast, and Fully Automatic Learning Algorithm for Predictive Coding Networks (ICLR 2024 · 8 citations)
normalization · bn · mini · 111 papers
Approaches in this cluster:
- Batch normalization improvements (42 papers)
Analyzes and improves batch normalization through decorrelation, stable statistics, and regularization insight. - Normalization layer theory (37 papers)
Explains generalization and nonlinearity of normalization layers and removes them with careful initialization. - Weight normalization and decay interplay (32 papers)
Studies weight decay and weight normalization dynamics, including periodic training behavior and sparse training skew.
Most cited and most cited since 2024:
- On Calibration of Modern Neural Networks (ICML 2017 · 2,519 citations)
- Self-Normalizing Neural Networks (NeurIPS 2017 · 511 citations)
- Batch Normalization Alleviates the Spectral Bias in Coordinate Networks (CVPR 2024 · 19 citations)
- Till the Layers Collapse: Compressing a Deep Neural Network Through the Lenses of Batch Normalization Layers. (AAAI 2025 · 3 citations)
Related topics in Architectures and efficiency
- Pruning and efficient tuning (1,147)
- Attention and transformers (2,296)
- Neural architecture search (380)
- Quantization (646)
- Spiking networks (289)
- Mixture of experts (300)
- CNN design (853)
