Video segmentation and tracking: research map
1,736 accepted papers on Video segmentation and tracking in Video, from ICML, NeurIPS, ICLR, CVPR and AAAI (2016–2026), grouped into 6 clusters and 21 approaches. The busiest year so far is 2026.
Within Video, its share shrank from 42.6% in 2023–24 to 21.3% in 2025–26 (440 → 537 papers at ICML, NeurIPS, CVPR and AAAI, the venues with data for all four years).
Explore Video segmentation and tracking in the interactive map
Working on something in this topic? Describe your idea in scime atlas to see which approach it falls under, the closest papers by meaning and how crowded the spot has become.
Approaches and key papers
detection · anomaly · supervised · 662 papers
Approaches in this cluster:
- Self-supervised video representation learning (167 papers)
Learns video features from temporal self-supervision, correspondence and reconstruction without labels. - Video restoration with temporal priors (159 papers)
Enhances and restores degraded video using internal priors, state-space models and temporal coherence. - Sign language and physiological video analysis (99 papers)
Applies temporal modeling to sign language recognition, visual speech and remote physiological measurement. - Weakly supervised video anomaly detection (99 papers)
Detects anomalous events in surveillance video using weak labels, memory units and event completeness. - Spatio-temporal predictive modeling (91 papers)
Models spatiotemporal sequences such as traffic using factorized, Mamba and attention-based architectures. - Deepfake video detection (47 papers)
Detects forged face videos via inconsistency learning, facial components and self-supervised real-face cues.
Most cited and most cited since 2024:
- Real-World Anomaly Detection in Surveillance Videos (CVPR 2018 · 2,159 citations)
- Celeb-DF: A Large-Scale Challenging Dataset for DeepFake Forensics (CVPR 2020 · 1,725 citations)
- Open-Vocabulary Video Anomaly Detection (CVPR 2024 · 90 citations)
- Exploiting Style Latent Flows for Generalizing Deepfake Video Detection (CVPR 2024 · 77 citations)
prediction · person · pose · 419 papers
Approaches in this cluster:
- Layered video decomposition and dynamics (140 papers)
Decomposes videos into layers, objects and effects to model scene dynamics and generate videos. - Stochastic video prediction (125 papers)
Predicts future frames with learned dynamics, hierarchical models and example guidance. - Physics-inspired motion prediction (113 papers)
Learns interaction dynamics and motion simulators from video for future motion forecasting. - Spatial-temporal video person re-identification (41 papers)
Matches people across video using spatiotemporal attention, graphs and keypoint message passing.
Most cited and most cited since 2024:
- BDD100K: A Diverse Driving Dataset for Heterogeneous Multitask Learning (CVPR 2020 · 2,481 citations)
- Face2Face: Real-Time Face Capture and Reenactment of RGB Videos (CVPR 2016 · 1,843 citations)
- ODTrack: Online Dense Temporal Token Learning for Visual Tracking (AAAI 2024 · 200 citations)
- Autoregressive Queries for Adaptive Tracking with Spatio-Temporal Transformers (CVPR 2024 · 160 citations)
interpolation · flow · event · 253 papers
Approaches in this cluster:
- Learned optical flow estimation (73 papers)
Estimates optical flow with multi-frame models, occlusion awareness and graph reasoning. - Video frame interpolation (60 papers)
Synthesizes intermediate frames using transformers, depth-aware warping and motion modeling. - Event-based video reconstruction (63 papers)
Reconstructs high-speed, high-dynamic-range video from event cameras with learned losses and simulation. - Spike and neuromorphic deblurring (57 papers)
Removes motion blur and noise using spike streams and neuromorphic event data.
Most cited and most cited since 2024:
- Super SloMo: High Quality Estimation of Multiple Intermediate Frames for Video Interpolation (CVPR 2018 · 867 citations)
- GMS: Grid-based Motion Statistics for Fast, Ultra-Robust Feature Correspondence (CVPR 2017 · 618 citations)
- HARDVS: Revisiting Human Activity Recognition with Dynamic Vision Sensors (AAAI 2024 · 48 citations)
- MemFlow: Optical Flow Estimation and Prediction with Memory (CVPR 2024 · 45 citations)
object segmentation · video object · vos · 220 papers
Approaches in this cluster:
- Mask propagation for unsupervised video segmentation (67 papers)
Segments objects in video with weak, unsupervised or training-free methods using object descriptors and correspondence. - Matching-based video object segmentation (86 papers)
Propagates object masks via feature matching, instance queries and efficient linear attention. - Video semantic and panoptic segmentation (67 papers)
Segments video frames through flow propagation, mask propagation and unified kernel baselines.
Most cited and most cited since 2024:
- A Benchmark Dataset and Evaluation Methodology for Video Object Segmentation (CVPR 2016 · 2,095 citations)
- One-Shot Video Object Segmentation (CVPR 2017 · 926 citations)
- MemSAM: Taming Segment Anything Model for Echocardiography Video Segmentation (CVPR 2024 · 62 citations)
- Decoupling Static and Hierarchical Motion Perception for Referring Video Segmentation (CVPR 2024 · 41 citations)
super resolution · video super · vsr · 93 papers
Approaches in this cluster:
- Real-world video super-resolution (60 papers)
Restores real-world low-resolution video with GAN detail synthesis, alignment and artifact mitigation. - Space-time video super-resolution (33 papers)
Jointly upsamples spatial and temporal resolution using trajectory-aware attention, events and dynamic filters.
Most cited and most cited since 2024:
- Real-Time Video Super-Resolution With Spatio-Temporal Networks and Motion Compensation (CVPR 2017 · 800 citations)
- Deep Video Super-Resolution Network Using Dynamic Upsampling Filters Without Explicit Motion Compensation (CVPR 2018 · 636 citations)
- Video Super-Resolution Transformer with Masked Inter&Intra-Frame Attention (CVPR 2024 · 36 citations)
- Enhancing Video Super-Resolution via Implicit Resampling-based Alignment (CVPR 2024 · 34 citations)
compression · coding · rate · 89 papers
Approaches in this cluster:
- Learned video compression (68 papers)
Replaces traditional codecs with end-to-end networks using motion compensation, feature-space prediction and transformers. - Neural video representations (21 papers)
Encodes videos as implicit networks with hybrid or frame-embedding designs for compression.
Most cited and most cited since 2024:
- DVC: An End-To-End Deep Video Compression Framework (CVPR 2019 · 744 citations)
- Scale-Space Flow for End-to-End Optimized Video Compression (CVPR 2020 · 316 citations)
- Neural Video Compression with Feature Modulation (CVPR 2024 · 140 citations)
- Towards Practical Real-Time Neural Video Compression (CVPR 2025 · 61 citations)
Related topics in Video
- Video-language understanding (1,146)
- Action recognition (717)
- Video generation (1,407)
