atlas

Research map / Vision-language and multimodal

Multimodal LLMs: research map

2,371 accepted papers on Multimodal LLMs in Vision-language and multimodal, from ICML, NeurIPS, ICLR, CVPR and AAAI (2016–2026), grouped into 4 clusters and 15 approaches. The busiest year so far is 2026.

Within Vision-language and multimodal, its share grew from 13.0% in 2023–24 to 45.1% in 2025–26 (207 → 2,156 papers at ICML, NeurIPS, CVPR and AAAI, the venues with data for all four years).

2016: 1162017: 0172018: 1182019: 1192020: 3202021: 2212022: 0222023: 8232024: 199242025: 658252026: 149826

Explore Multimodal LLMs in the interactive map

Working on something in this topic? Describe your idea in scime atlas to see which approach it falls under, the closest papers by meaning and how crowded the spot has become.

Approaches and key papers

vlms · visual reasoning · chain · 1,218 papers

Approaches in this cluster:

Most cited and most cited since 2024:

mllm · instruction · llms · 753 papers

Approaches in this cluster:

Most cited and most cited since 2024:

tokens · pruning · visual token · 207 papers

Approaches in this cluster:

Most cited and most cited since 2024:

hallucination · lvlms · mitigating · 193 papers

Approaches in this cluster:

Most cited and most cited since 2024:

Related topics in Vision-language and multimodal