Research map / Architectures and efficiency
Quantization: research map
646 accepted papers on Quantization in Architectures and efficiency, from ICML, NeurIPS, ICLR, CVPR and AAAI (2016–2026), grouped into 2 clusters and 8 approaches. The busiest year so far is 2026.
Within Architectures and efficiency, its share grew from 6.0% in 2023–24 to 9.6% in 2025–26 (118 → 341 papers at ICML, NeurIPS, CVPR and AAAI, the venues with data for all four years).
Explore Quantization in the interactive map
Working on something in this topic? Describe your idea in scime atlas to see which approach it falls under, the closest papers by meaning and how crowded the spot has become.
Approaches and key papers
post quantization · ptq · compression · 413 papers
Approaches in this cluster:
- Post-training weight quantization (100 papers)
Compresses weights by quantization with sharpness control, loss-aware objectives, and quantization noise. - Post-training quantization of LLMs (136 papers)
Handles outliers and low-bit errors when quantizing large language models after training. - Post-training quantization of ViTs (105 papers)
Reduces quantization error in vision transformers using Hessian, Fisher, and outlier-aware reconstruction. - KV cache quantization (72 papers)
Compresses key-value caches with vector and mixed-precision quantization for long-context inference.
Most cited and most cited since 2024:
- QLoRA: Efficient Finetuning of Quantized LLMs (NeurIPS 2023 · 824 citations)
- Soft-to-Hard Vector Quantization for End-to-End Learning Compressible Representations (NeurIPS 2017 · 358 citations)
- OWQ: Outlier-Aware Weight Quantization for Efficient Fine-Tuning and Inference of Large Language Models (AAAI 2024 · 72 citations)
- KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization (NeurIPS 2024 · 51 citations)
mixed precision · floating · point · 233 papers
Approaches in this cluster:
- Mixed-precision quantization (92 papers)
Assigns bit-widths per layer for low-bit network training and inference with learnable quantizers. - Low-precision training formats (58 papers)
Designs 8-bit floating-point formats and accumulation schemes for low-precision neural network training. - Ultra-low-bit LLM quantization (56 papers)
Pushes language models to FP4, 1-bit, and multi-bit formats with scaling-law analysis. - Binary neural networks (27 papers)
Binarizes weights and activations with information-retention and ensemble techniques.
Most cited and most cited since 2024:
- Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference (CVPR 2018 · 3,940 citations)
- Mixed Precision Training (ICLR 2018 · 999 citations)
- JointSQ: Joint Sparsification-Quantization for Distributed Learning (CVPR 2024 · 18 citations)
- Towards Efficient Verification of Quantized Neural Networks (AAAI 2024 · 17 citations)
Related topics in Architectures and efficiency
- Pruning and efficient tuning (1,147)
- Attention and transformers (2,296)
- Neural architecture search (380)
- RNNs and normalization (2,827)
- Spiking networks (289)
- Mixture of experts (300)
- CNN design (853)
