Human pose estimation: research map
1,051 accepted papers on Human pose estimation in 3D vision, from ICML, NeurIPS, ICLR, CVPR and AAAI (2016–2026), grouped into 3 clusters and 14 approaches. The busiest year so far is 2026.
Within 3D vision, its share shrank from 15.2% in 2023–24 to 10.8% in 2025–26 (271 → 300 papers at ICML, NeurIPS, CVPR and AAAI, the venues with data for all four years).
Explore Human pose estimation in the interactive map
Working on something in this topic? Describe your idea in scime atlas to see which approach it falls under, the closest papers by meaning and how crowded the spot has become.
Approaches and key papers
object pose · 6d · hand · 400 papers
Approaches in this cluster:
- Depth-based hand pose regression (84 papers)
Estimates 3D hand pose from depth or RGB images using CNN regression, dense prediction and neural rendering. - Learned camera pose estimation (96 papers)
Estimates absolute and relative camera pose using learned features, pose regression and geometric solvers. - Category-level object pose estimation (92 papers)
Estimates object pose for unseen instances within categories using correspondence learning and universal features. - Instance-level 6D object pose (74 papers)
Predicts 6D poses of known or novel objects via iterative refinement, viewpoint encoding and few-shot matching. - Category-level pose estimation (54 papers)
Estimate object pose and part motion from point clouds.
Most cited and most cited since 2024:
- DenseFusion: 6D Object Pose Estimation by Iterative Dense Fusion (CVPR 2019 · 1,158 citations)
- PVNet: Pixel-Wise Voting Network for 6DoF Pose Estimation (CVPR 2019 · 1,079 citations)
- FoundationPose: Unified 6D Pose Estimation and Tracking of Novel Objects (CVPR 2024 · 315 citations)
- Reconstructing Hands in 3D with Transformers (CVPR 2024 · 151 citations)
human pose · pose estimation · 3d human · 350 papers
Approaches in this cluster:
- Monocular 2D-to-3D pose lifting (90 papers)
Lifts single-image or video 2D keypoints to 3D with multi-hypothesis, weak-supervision and adaptation methods. - Transformer-based multi-view 3D pose (83 papers)
Uses transformers and diffusion with multi-view or occlusion-aware refinement for 3D human pose. - Multi-person pose estimation (51 papers)
Detects and assigns keypoints for many people using top-down, bottom-up or single-stage approaches. - Heatmap and coordinate pose encoding (50 papers)
Refines keypoint encoding, decoding and probabilistic output representations for 2D human pose. - Temporal video pose estimation (76 papers)
Estimates human pose in video using temporal difference learning and sparse labels.
Most cited and most cited since 2024:
- Realtime Multi-Person 2D Pose Estimation Using Part Affinity Fields (CVPR 2017 · 7,514 citations)
- Deep High-Resolution Representation Learning for Human Pose Estimation (CVPR 2019 · 5,835 citations)
- KTPFormer: Kinematics and Trajectory Prior Knowledge-Enhanced Transformer for 3D Human Pose Estimation (CVPR 2024 · 107 citations)
- RTMO: Towards High-Performance One-Stage Real-Time Multi-Person Pose Estimation (CVPR 2024 · 97 citations)
body · motion · human mesh · 301 papers
Approaches in this cluster:
- Physics and sensor-based motion capture (125 papers)
Recovers human motion from video or inertial sensors using physical constraints and object interaction. - Human shape and pose regression (80 papers)
Estimates body shape and pose with mesh regression, probabilistic and whole-body models. - Human mesh recovery from images (58 papers)
Recovers 3D body meshes end-to-end with pose calibration, point guidance and diffusion. - Human motion prediction (38 papers)
Forecasts future 3D human motion from past observations using dynamic relationship modeling.
Most cited and most cited since 2024:
- End-to-End Recovery of Human Shape and Pose (CVPR 2018 · 1,958 citations)
- Expressive Body Capture: 3D Hands, Face, and Body From a Single Image (CVPR 2019 · 1,796 citations)
- WHAM: Reconstructing World-grounded Humans with Accurate 3D Motion (CVPR 2024 · 98 citations)
- TokenHMR: Advancing Human Mesh Recovery with a Tokenized Pose Representation (CVPR 2024 · 70 citations)
Related topics in 3D vision
- Gaussian splatting (841)
- NeRF and radiance fields (456)
- 3D generation (1,606)
- Face and shape (1,286)
- Depth and stereo (1,365)
