Depth and stereo: research map
1,365 accepted papers on Depth and stereo in 3D vision, from ICML, NeurIPS, ICLR, CVPR and AAAI (2016–2026), grouped into 4 clusters and 16 approaches. The busiest year so far is 2026.
Within 3D vision, its share shrank from 17.2% in 2023–24 to 14.8% in 2025–26 (306 → 409 papers at ICML, NeurIPS, CVPR and AAAI, the venues with data for all four years).
Explore Depth and stereo in the interactive map
Working on something in this topic? Describe your idea in scime atlas to see which approach it falls under, the closest papers by meaning and how crowded the spot has become.
Approaches and key papers
motion · flow · localization · 532 papers
Approaches in this cluster:
- Feed-forward 3D reconstruction (133 papers)
Reconstructs scenes and structure-from-motion in a single pass with learned priors, planes and dense matching. - Camera model and calibration for SfM (99 papers)
Handles camera calibration, distortion and minimal geometric solvers in large-scale structure from motion. - Learned optical and scene flow (108 papers)
Estimates motion fields with unrolled, unsupervised or joint learning of flow, depth and ego-motion. - Scene coordinate regression for localization (112 papers)
Localizes cameras by regressing scene coordinates or maps, with voting and 3D surface prediction. - Dense feature matching for correspondence (80 papers)
Learns dense or two-view correspondences with kernelized matching, cycle consistency and confidence estimation.
Most cited and most cited since 2024:
- Structure-From-Motion Revisited (CVPR 2016 · 6,291 citations)
- SuperGlue: Learning Feature Matching With Graph Neural Networks (CVPR 2020 · 2,845 citations)
- DUSt3R: Geometric 3D Vision Made Easy (CVPR 2024 · 545 citations)
- SplaTAM: Splat Track & Map 3D Gaussians for Dense RGB-D SLAM (CVPR 2024 · 438 citations)
depth estimation · monocular · depth map · 342 papers
Approaches in this cluster:
- Self-supervised monocular depth (116 papers)
Learns depth from unlabeled images or video using geometric constraints, semantic guidance and planarity priors. - Dense depth prediction and completion (85 papers)
Predicts consistent, high-resolution or sparse-input-completed depth from single images and video. - Foundation models for monocular geometry (81 papers)
Estimates geometry from single images with inductive biases, open-domain training and synthetic-real adaptation. - Depth from light field and defocus (60 papers)
Estimates depth using light fields, defocus, event cues and interference instead of standard RGB.
Most cited and most cited since 2024:
- Unsupervised Monocular Depth Estimation With Left-Right Consistency (CVPR 2017 · 3,325 citations)
- Unsupervised Learning of Depth and Ego-Motion From Video (CVPR 2017 · 2,885 citations)
- VGGT: Visual Geometry Grounded Transformer (CVPR 2025 · 410 citations)
- Repurposing Diffusion-Based Image Generators for Monocular Depth Estimation (CVPR 2024 · 388 citations)
light · imaging · illumination · 288 papers
Approaches in this cluster:
- Neural light field and illumination imaging (107 papers)
Recovers scene properties with learned light fields, structured light and material estimation. - Non-line-of-sight imaging (93 papers)
Reconstructs hidden or transparent geometry from time-of-flight light transport. - Spike and event camera reconstruction (50 papers)
Reconstructs high-speed scenes from spike, single-photon and neuromorphic sensor data. - Snapshot hyperspectral imaging (38 papers)
Reconstructs spectral images from coded or dispersive snapshots with learned or physics-guided models.
Most cited and most cited since 2024:
- DeepView: View Synthesis With Learned Gradient Descent (CVPR 2019 · 434 citations)
- EPINET: A Fully-Convolutional Neural Network Using Epipolar Geometry for Depth From Light Field Images (CVPR 2018 · 317 citations)
- Boosting Spike Camera Image Reconstruction from a Perspective of Dealing with Spike Fluctuations (CVPR 2024 · 16 citations)
- Progressive Divide-and-Conquer via Subsampling Decomposition for Accelerated MRI (CVPR 2024 · 16 citations)
view stereo · photometric stereo · stereo matching · 203 papers
Approaches in this cluster:
- Cost-volume multi-view stereo (107 papers)
Infers depth from multiple views with learned cost volumes, recurrent regularization and geometry-aware networks. - Photometric stereo (43 papers)
Estimates surface normals from varying illumination, including uncalibrated and non-Lambertian cases. - Learned stereo with specialized sensors (53 papers)
Estimates depth from stereo pairs with active, polarization, normal or exposure cues.
Most cited and most cited since 2024:
- Pyramid Stereo Matching Network (CVPR 2018 · 1,860 citations)
- A Multi-View Stereo Benchmark With High-Resolution Images and Multi-Camera Videos (CVPR 2017 · 913 citations)
- MoCha-Stereo: Motif Channel Attention Network for Stereo Matching (CVPR 2024 · 86 citations)
- MonSter: Marry Monodepth to Stereo Unleashes Power (CVPR 2025 · 53 citations)
Related topics in 3D vision
- Gaussian splatting (841)
- Human pose estimation (1,051)
- NeRF and radiance fields (456)
- 3D generation (1,606)
- Face and shape (1,286)
