Video generation: research map
1,407 accepted papers on Video generation in Video, from ICML, NeurIPS, ICLR, CVPR and AAAI (2016–2026), grouped into 2 clusters and 10 approaches. The busiest year so far is 2026.
Within Video, its share grew from 22.1% in 2023–24 to 44.5% in 2025–26 (228 → 1,121 papers at ICML, NeurIPS, CVPR and AAAI, the venues with data for all four years).
Explore Video generation in the interactive map
Working on something in this topic? Describe your idea in scime atlas to see which approach it falls under, the closest papers by meaning and how crowded the spot has become.
Approaches and key papers
video generation · video diffusion · text video · 892 papers
Approaches in this cluster:
- Efficient video diffusion models (283 papers)
Builds video diffusion models with latent decomposition, autoregressive distillation and training-free long-video generation. - Compositional text-to-video diffusion (159 papers)
Generates videos from text with improved motion, composition and extendable length, plus benchmarks. - Trajectory-guided motion control (184 papers)
Controls video generation motion with sparse trajectories, latent trajectory guidance and tracking signals. - Autoregressive long video generation (156 papers)
Generates long videos autoregressively with continuous tokens, propagation and rich perception, plus evaluation suites. - World-model video generation (110 papers)
Uses video generators as interactive world simulators and benchmarks their physical consistency.
Most cited and most cited since 2024:
- MoCoGAN: Decomposing Motion and Content for Video Generation (CVPR 2018 · 1,116 citations)
- VBench: Comprehensive Benchmark Suite for Video Generative Models (CVPR 2024 · 351 citations)
- VideoCrafter2: Overcoming Data Limitations for High-Quality Video Diffusion Models (CVPR 2024 · 223 citations)
- Follow Your Pose: Pose-Guided Text-to-Video Generation Using Pose-Free Videos (AAAI 2024 · 122 citations)
audio · editing · talking · 515 papers
Approaches in this cluster:
- Diffusion-based human image animation (126 papers)
Animates portraits and human images with video diffusion, controllable motion and camera conditions. - Diffusion-based video editing (92 papers)
Edits video consistently with diffusion feature correspondence, reward tuning and frequency-aware factorization. - Talking face generation (81 papers)
Generates audio-driven talking faces from single images with landmark priors and motion diffusion. - Video-to-audio synthesis (76 papers)
Generates synchronized soundtracks from video with latent diffusion and text conditioning. - Text-to-motion diffusion (140 papers)
Synthesize human motion from text using diffusion models.
Most cited and most cited since 2024:
- One-Shot Free-View Neural Talking-Head Synthesis for Video Conferencing (CVPR 2021 · 445 citations)
- Hierarchical Cross-Modal Talking Face Generation With Dynamic Pixel-Wise Loss (CVPR 2019 · 440 citations)
- MagicAnimate: Temporally Consistent Human Image Animation using Diffusion Model (CVPR 2024 · 143 citations)
- Video-P2P: Video Editing with Cross-attention Control (CVPR 2024 · 123 citations)
Related topics in Video
- Video segmentation and tracking (1,736)
- Video-language understanding (1,146)
- Action recognition (717)
