3D generation: research map
1,606 accepted papers on 3D generation in 3D vision, from ICML, NeurIPS, ICLR, CVPR and AAAI (2016–2026), grouped into 4 clusters and 16 approaches. The busiest year so far is 2026.
Within 3D vision, its share grew from 25.4% in 2023–24 to 31.9% in 2025–26 (453 → 881 papers at ICML, NeurIPS, CVPR and AAAI, the venues with data for all four years).
Explore 3D generation in the interactive map
Working on something in this topic? Describe your idea in scime atlas to see which approach it falls under, the closest papers by meaning and how crowded the spot has become.
Approaches and key papers
semantic · occupancy · driving · 566 papers
Approaches in this cluster:
- Learning 3D reconstruction from images (196 papers)
Reconstructs 3D structure from single or multiple views with learned priors and foundation-model features. - Semantic scene completion (153 papers)
Predicts complete semantic 3D scenes from monocular images using foundation-model guidance and scene representations. - 2D-to-3D pretrained representations (94 papers)
Transfers pretrained 2D models to 3D, medical and molecular representation learning. - 3D perception for autonomous driving (73 papers)
Targets lanes, mapping, pretraining and rendering for driving, with large benchmarks. - Vision-based 3D occupancy prediction (50 papers)
Predicts semantic occupancy with efficient representations such as octrees, superquadrics and low-rank decomposition.
Most cited and most cited since 2024:
- Scalability in Perception for Autonomous Driving: Waymo Open Dataset (CVPR 2020 · 3,252 citations)
- Volumetric and Multi-View CNNs for Object Classification on 3D Data (CVPR 2016 · 1,610 citations)
- Reliable Conflictive Multi-View Learning (AAAI 2024 · 126 citations)
- Continuous 3D Perception Model with Persistent State (CVPR 2025 · 81 citations)
3d generation · text · consistent · 436 papers
Approaches in this cluster:
- Diffusion-based novel view synthesis (156 papers)
Generates consistent novel views from one or few images using multi-view and video diffusion priors. - Score distillation for text-to-3D (106 papers)
Distills 2D diffusion priors into 3D with multi-view consistency to produce text-conditioned assets. - Feed-forward 3D asset generation and editing (116 papers)
Generates, textures and edits 3D content using reconstruction models and decomposed outputs. - Text-to-3D scene generation (58 papers)
Creates 3D scenes from text via latent diffusion, direct training and compositional factoring.
Most cited and most cited since 2024:
- Zero-Shot Text-Guided Object Generation With Dream Fields (CVPR 2022 · 381 citations)
- Wonder3D: Single Image to 3D using Cross-Domain Diffusion (CVPR 2024 · 341 citations)
- Structured 3D Latents for Scalable and Versatile 3D Generation (CVPR 2025 · 247 citations)
- One-2-3-45++: Fast Single Image to 3D Objects with Consistent Multi-View Generation and 3D Diffusion (CVPR 2024 · 141 citations)
motion · interaction · human · 366 papers
Approaches in this cluster:
- Monocular 3D object reconstruction (94 papers)
Reconstructs objects and dynamic scenes from monocular input using synthetic data and pose-free online methods. - Hand-object interaction generation (116 papers)
Generates 3D hand-object and grasp motion with contact guidance and diffusion models. - Scene-aware human motion synthesis (82 papers)
Generates or predicts human motion conditioned on 3D scenes, text and gaze. - Learning physics from video (74 papers)
Learns simulators and physical properties from videos, often with Gaussian-based scene representations.
Most cited and most cited since 2024:
- Hand Keypoint Detection in Single Images Using Multiview Bootstrapping (CVPR 2017 · 1,182 citations)
- Generating Diverse and Natural 3D Human Motions From Text (CVPR 2022 · 553 citations)
- EMAGE: Towards Unified Holistic Co-Speech Gesture Generation via Expressive Masked Audio Gesture Modeling (CVPR 2024 · 79 citations)
- Editable Scene Simulation for Autonomous Driving via Collaborative LLM-Agents (CVPR 2024 · 73 citations)
layout · indoor · 3d scene · 238 papers
Approaches in this cluster:
- Generative indoor scene synthesis (131 papers)
Generates 3D indoor layouts using scene graphs, diffusion, grammars and language-driven spatial reasoning. - Open-vocabulary 3D scene understanding (95 papers)
Aligns 3D scenes with language for open-world, part-level and affordance understanding. - Visual-acoustic scene modeling (12 papers)
Synthesizes room acoustics and novel-view sound using simulation and learned acoustic fields.
Most cited and most cited since 2024:
- ScanNet: Richly-Annotated 3D Reconstructions of Indoor Scenes (CVPR 2017 · 4,275 citations)
- OpenScene: 3D Scene Understanding With Open Vocabularies (CVPR 2023 · 365 citations)
- DiffuScene: Denoising Diffusion Models for Generative Indoor Scene Synthesis (CVPR 2024 · 93 citations)
- Holodeck: Language Guided Generation of 3D Embodied AI Environments (CVPR 2024 · 91 citations)
Related topics in 3D vision
- Gaussian splatting (841)
- Human pose estimation (1,051)
- NeRF and radiance fields (456)
- Face and shape (1,286)
- Depth and stereo (1,365)
