Research map / Image generation
Autoregressive and latent generation: research map
1,747 accepted papers on Autoregressive and latent generation in Image generation, from ICML, NeurIPS, ICLR, CVPR and AAAI (2016–2026), grouped into 6 clusters and 21 approaches. The busiest year so far is 2026.
Within Image generation, its share held steady from 22.7% in 2023–24 to 22.0% in 2025–26 (474 → 901 papers at ICML, NeurIPS, CVPR and AAAI, the venues with data for all four years).
Explore Autoregressive and latent generation in the interactive map
Working on something in this topic? Describe your idea in scime atlas to see which approach it falls under, the closest papers by meaning and how crowded the spot has become.
Approaches and key papers
latent space · modeling · vae · 557 papers
Approaches in this cluster:
- Generative latent space tooling (129 papers)
Reuse and analyze latent spaces and tokenizers of image generators. - High-resolution diffusion synthesis (140 papers)
Scale diffusion models to high resolutions with position extrapolation and efficient architectures. - Latent diffusion with learned tokenizers (114 papers)
Apply diffusion and masked modeling in learned latent spaces for images and language. - Hierarchical VQ-VAE generation (62 papers)
Generate images with hierarchical quantized or multi-scale latent representations. - Next-scale and token autoregressive images (90 papers)
Generate images autoregressively over scales or continuous tokens with improved tokenizers. - Accelerating autoregressive image generation (22 papers)
Speed autoregressive and masked image generators with speculative decoding and hierarchical modeling.
Most cited and most cited since 2024:
- High-Resolution Image Synthesis With Latent Diffusion Models (CVPR 2022 · 15,775 citations)
- Taming Transformers for High-Resolution Image Synthesis (CVPR 2021 · 2,386 citations)
- Efficient Deformable ConvNets: Rethinking Dynamic and Sparse Operator for Vision Applications (CVPR 2024 · 257 citations)
- Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction (NeurIPS 2024 · 111 citations)
captioning · assessment · synthetic · 460 papers
Approaches in this cluster:
- Classical generative image synthesis (147 papers)
Synthesize and inpaint images with CNN, GAN, and example-based generators. - Retrieval and matching with generative features (140 papers)
Apply generative or diffusion features to re-ranking, matching, and metric learning. - Synthetic training data from generators (111 papers)
Use generated images for classification training and evaluate generative models. - Learning-based image quality assessment (62 papers)
Predict perceptual image quality without reference using learned or synthetic supervision.
Most cited and most cited since 2024:
- Generative Image Inpainting With Contextual Attention (CVPR 2018 · 2,538 citations)
- BERTScore: Evaluating Text Generation with BERT (ICLR 2020 · 2,397 citations)
- U-KAN Makes Strong Backbone for Medical Image Segmentation and Generation (AAAI 2025 · 308 citations)
- MedSegDiff-V2: Diffusion-Based Medical Image Segmentation with Transformer (AAAI 2024 · 263 citations)
ai · watermarking · generated · 243 papers
Approaches in this cluster:
- Generalizable fake image detection (119 papers)
Detect AI-generated images with large-scale datasets, multi-generator training, and semantic-pixel features. - Robust diffusion watermarking (79 papers)
Embed and attack watermarks in diffusion-generated images for provenance. - Reconstruction-based diffusion image detection (45 papers)
Detect diffusion-generated images using reconstruction error and diffusion-based anomaly modeling.
Most cited and most cited since 2024:
- CNN-Generated Images Are Surprisingly Easy to Spot... for Now (CVPR 2020 · 1,121 citations)
- Towards Universal Fake Image Detectors That Generalize Across Generative Models (CVPR 2023 · 381 citations)
- Forgery-aware Adaptive Transformer for Generalizable Synthetic Image Detection (CVPR 2024 · 112 citations)
- EditGuard: Versatile Image Watermarking for Tamper Localization and Copyright Protection (CVPR 2024 · 91 citations)
motion · layout · object · 202 papers
Approaches in this cluster:
- Compositional scene generation (105 papers)
Generate images from scene graphs and compositional structure. - Diffusion planning and policies (58 papers)
Use diffusion models for robot planning, behavior synthesis, and visuomotor control. - Layout generation (39 papers)
Generate graphic and design layouts with transformers and diffusion models.
Most cited and most cited since 2024:
- Image Generation From Scene Graphs (CVPR 2018 · 854 citations)
- Executing Your Commands via Motion Diffusion in Latent Space (CVPR 2023 · 341 citations)
- DisCo: Disentangled Control for Realistic Human Dance Generation (CVPR 2024 · 79 citations)
- SingularTrajectory: Universal Trajectory Predictor Using Diffusion Model (CVPR 2024 · 67 citations)
speech · music · audio · 166 papers
Approaches in this cluster:
- Neural text-to-speech synthesis (72 papers)
Synthesize speech from text with codec, diffusion, and vector-quantized generative models. - Diffusion audio synthesis (52 papers)
Generate waveforms and general audio with diffusion and latent diffusion models. - Controllable music generation (42 papers)
Generate symbolic and audio music with hierarchical, autoregressive, and diffusion models.
Most cited and most cited since 2024:
- FastSpeech: Fast, Robust and Controllable Text to Speech (NeurIPS 2019 · 579 citations)
- SampleRNN: An Unconditional End-to-End Neural Audio Generation Model (ICLR 2017 · 423 citations)
- DiffSHEG: A Diffusion-Based Approach for Real-Time Speech-driven Holistic 3D Expression and Gesture Generation (CVPR 2024 · 49 citations)
- NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers (ICLR 2024 · 37 citations)
protein · molecular · molecule · 119 papers
Approaches in this cluster:
- Equivariant diffusion for 3D molecules (71 papers)
Generate 3D molecular structures with equivariant diffusion in geometric or latent space. - Diffusion for protein and antibody design (48 papers)
Generate protein structures, sequences, and antibodies using latent and language-model diffusion.
Most cited and most cited since 2024:
- Antigen-Specific Antibody Design and Optimization with Diffusion-Based Generative Models for Protein Structures (NeurIPS 2022 · 119 citations)
- Equivariant Diffusion for Molecule Generation in 3D (ICML 2022 · 118 citations)
- Deep Confident Steps to New Pockets: Strategies for Docking Generalization (ICLR 2024 · 28 citations)
- Scalable Diffusion for Materials Generation (ICLR 2024 · 25 citations)
Related topics in Image generation
- GANs (515)
- Super-resolution (429)
- Editing and style transfer (988)
- Fast diffusion sampling (1,866)
- Image restoration (1,184)
- Text-to-image diffusion (1,310)
