Research map / Vision-language and multimodal
Visual question answering: research map
546 accepted papers on Visual question answering in Vision-language and multimodal, from ICML, NeurIPS, ICLR, CVPR and AAAI (2016–2026), grouped into 2 clusters and 7 approaches. The busiest year so far is 2025.
Within Vision-language and multimodal, its share shrank from 7.0% in 2023–24 to 3.6% in 2025–26 (111 → 173 papers at ICML, NeurIPS, CVPR and AAAI, the venues with data for all four years).
Explore Visual question answering in the interactive map
Working on something in this topic? Describe your idea in scime atlas to see which approach it falls under, the closest papers by meaning and how crowded the spot has become.
Approaches and key papers
visual question · answering vqa · knowledge · 321 papers
Approaches in this cluster:
- Universal VQA models (100 papers)
Develops general and knowledge-based visual question answering with memory, reasoning transfer, and causal adapters. - Attention and prior-robust VQA (109 papers)
Improves VQA via multi-level attention, bias reduction, decomposition, and cycle consistency. - Vision-language model VQA evaluation (56 papers)
Generates challenging benchmarks and zero-shot prompting for evaluating vision-language models on VQA. - Multimodal LLM question answering (56 papers)
Augments multimodal LLMs for knowledge-based and text-rich VQA and evaluates them dynamically.
Most cited and most cited since 2024:
- Bottom-Up and Top-Down Attention for Image Captioning and Visual Question Answering (CVPR 2018 · 5,159 citations)
- Making the v in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering (CVPR 2017 · 2,316 citations)
- NuScenes-QA: A Multi-Modal Visual Question Answering Benchmark for Autonomous Driving Scenario (AAAI 2024 · 157 citations)
- Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models (CVPR 2024 · 120 citations)
answer · qa · video question · 225 papers
Approaches in this cluster:
- Multimodal question answering datasets (90 papers)
Creates benchmarks and models for multi-hop, video, and text-based multimodal question answering. - Visual reasoning diagnostics (62 papers)
Diagnoses shortcuts and compositional reasoning in visual reasoning models using synthetic benchmarks. - Open-domain multimodal QA (73 papers)
Answers questions over multiple evidence sources with co-attention, cross-lingual retrieval, and dialog.
Most cited and most cited since 2024:
- Stacked Attention Networks for Image Question Answering (CVPR 2016 · 2,094 citations)
- CLEVR: A Diagnostic Dataset for Compositional Language and Elementary Visual Reasoning (CVPR 2017 · 1,831 citations)
- Generative Multimodal Models are In-Context Learners (CVPR 2024 · 97 citations)
- Can I Trust Your Answer? Visually Grounded Video Question Answering (CVPR 2024 · 62 citations)
Related topics in Vision-language and multimodal
- Human and embodied modeling (1,673)
- Multimodal LLMs (2,371)
- Robot action and navigation (630)
- CLIP and zero-shot (959)
- Cross-modal learning (1,303)
