Video-language understanding: research map
1,146 accepted papers on Video-language understanding in Video, from ICML, NeurIPS, ICLR, CVPR and AAAI (2016–2026), grouped into 3 clusters and 11 approaches. The busiest year so far is 2026.
Within Video, its share grew from 21.6% in 2023–24 to 28.9% in 2025–26 (223 → 728 papers at ICML, NeurIPS, CVPR and AAAI, the venues with data for all four years).
Explore Video-language understanding in the interactive map
Working on something in this topic? Describe your idea in scime atlas to see which approach it falls under, the closest papers by meaning and how crowded the spot has become.
Approaches and key papers
video understanding · question · long · 759 papers
Approaches in this cluster:
- Event and anomaly reasoning in video (184 papers)
Reasons about events, anomalies and object state changes in video with datasets and grounded reasoning. - Video LLMs with fine-grained understanding (229 papers)
Builds video language models for object-level and temporal understanding, with memory and benchmarks. - Long video understanding (236 papers)
Handles hour-long videos using memory, key-frame reasoning and training-free or streaming approaches. - Video question answering (64 papers)
Answers questions about videos using memory prompting, sampling, tool use and compositional benchmarks. - Egocentric video understanding (46 papers)
Benchmarks and trains multimodal models for first-person video reasoning and grounded question answering.
Most cited and most cited since 2024:
- MovieQA: Understanding Stories in Movies Through Question-Answering (CVPR 2016 · 681 citations)
- Long-Term Feature Banks for Detailed Video Understanding (CVPR 2019 · 461 citations)
- MVBench: A Comprehensive Multi-modal Video Understanding Benchmark (CVPR 2024 · 241 citations)
- MovieChat: From Dense Token to Sparse Memory for Long Video Understanding (CVPR 2024 · 154 citations)
grounding · captioning · sentence · 205 papers
Approaches in this cluster:
- Temporal video grounding (88 papers)
Localizes sentences in videos using local-global interaction, metric learning and unsupervised or scalable grounding. - Encoder-decoder video captioning (52 papers)
Generates video captions with hierarchical, object-aware and reinforcement-learned encoders. - Dense video captioning (65 papers)
Localizes and describes multiple events using joint modeling, memory retrieval and weak supervision.
Most cited and most cited since 2024:
- Lip Reading Sentences in the Wild (CVPR 2017 · 613 citations)
- End-to-End Dense Video Captioning With Masked Transformer (CVPR 2018 · 578 citations)
- VTimeLLM: Empower LLM to Grasp Video Moments (CVPR 2024 · 119 citations)
- ShareGPT4Video: Improving Video Understanding and Generation with Better Captions (NeurIPS 2024 · 56 citations)
retrieval · text · moment · 182 papers
Approaches in this cluster:
- Video-language pretraining (82 papers)
Pretrains video-text models with structured alignment, distillation and large-scale transcript data. - Text-video retrieval (59 papers)
Retrieves videos from text using coarse-to-fine, probabilistic and ambiguity-aware representations. - Video moment retrieval (41 papers)
Localizes text-described moments in videos with clip trimming, unified retrieval-highlight and generative approaches.
Most cited and most cited since 2024:
- MSR-VTT: A Large Video Description Dataset for Bridging Video and Language (CVPR 2016 · 1,822 citations)
- Less Is More: ClipBERT for Video-and-Language Learning via Sparse Sampling (CVPR 2021 · 577 citations)
- Text Prompt with Normality Guidance for Weakly Supervised Video Anomaly Detection (CVPR 2024 · 95 citations)
- Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers (CVPR 2024 · 81 citations)
Related topics in Video
- Video segmentation and tracking (1,736)
- Action recognition (717)
- Video generation (1,407)
