Action recognition: research map
717 accepted papers on Action recognition in Video, from ICML, NeurIPS, ICLR, CVPR and AAAI (2016–2026), grouped into 2 clusters and 8 approaches. The busiest year so far is 2022.
Within Video, its share shrank from 13.7% in 2023–24 to 5.2% in 2025–26 (141 → 132 papers at ICML, NeurIPS, CVPR and AAAI, the venues with data for all four years).
Explore Action recognition in the interactive map
Working on something in this topic? Describe your idea in scime atlas to see which approach it falls under, the closest papers by meaning and how crowded the spot has become.
Approaches and key papers
action recognition · video recognition · representation · 384 papers
Approaches in this cluster:
- Disentangled action recognition (182 papers)
Recognizes actions by decomposing dynamics, components and human-object interactions. - Spatiotemporal CNNs for action recognition (82 papers)
Designs 3D and residual convolutional architectures for video action recognition. - Contrastive video representation learning (74 papers)
Learns video features with contrastive objectives, continuity, motion decoupling and language concepts. - Skeleton-based action recognition (46 papers)
Recognizes actions from skeleton sequences with graph convolutions, LSTMs and prototypes.
Most cited and most cited since 2024:
- Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset (CVPR 2017 · 9,707 citations)
- A Closer Look at Spatiotemporal Convolutions for Action Recognition (CVPR 2018 · 3,632 citations)
- Revealing Key Details to See Differences: A Novel Prototypical Perspective for Skeleton-based Action Recognition (CVPR 2025 · 36 citations)
- ExACT: Language-guided Conceptual Reasoning and Uncertainty Estimation for Event-based Action Recognition and More (CVPR 2024 · 24 citations)
temporal action · action localization · action detection · 333 papers
Approaches in this cluster:
- Temporal action segmentation (88 papers)
Segments activities in untrimmed videos with unsupervised, weakly or semi-supervised clustering and transformers. - Weakly supervised temporal action localization (79 papers)
Localizes actions using video-level labels with snippet modeling, uncertainty and graph networks. - Fully supervised temporal action detection (87 papers)
Detects action boundaries using temporal convolution, multi-task recycling and context aggregation. - Action anticipation and latent actions (79 papers)
Predicts future actions with LLMs, graph growth and latent action world models.
Most cited and most cited since 2024:
- Temporal Convolutional Networks for Action Segmentation and Detection (CVPR 2017 · 2,175 citations)
- AVA: A Video Dataset of Spatio-Temporally Localized Atomic Visual Actions (CVPR 2018 · 1,021 citations)
- End-to-End Temporal Action Detection with 1B Parameters Across 1000 Frames (CVPR 2024 · 53 citations)
- FACT: Frame-Action Cross-Attention Temporal Modeling for Efficient Action Segmentation (CVPR 2024 · 52 citations)
Related topics in Video
- Video segmentation and tracking (1,736)
- Video-language understanding (1,146)
- Video generation (1,407)
