atlas

Research map / Vision-language and multimodal

Human and embodied modeling: research map

1,673 accepted papers on Human and embodied modeling in Vision-language and multimodal, from ICML, NeurIPS, ICLR, CVPR and AAAI (2016–2026), grouped into 5 clusters and 21 approaches. The busiest year so far is 2025.

Within Vision-language and multimodal, its share shrank from 33.5% in 2023–24 to 14.0% in 2025–26 (532 → 669 papers at ICML, NeurIPS, CVPR and AAAI, the venues with data for all four years).

2016: 29162017: 45172018: 43182019: 59192020: 67202021: 109212022: 120222023: 213232024: 319242025: 377252026: 29226

Explore Human and embodied modeling in the interactive map

Working on something in this topic? Describe your idea in scime atlas to see which approach it falls under, the closest papers by meaning and how crowded the spot has become.

Approaches and key papers

retrieval · scene · captioning · 554 papers

Approaches in this cluster:

Most cited and most cited since 2024:

pre · generation · unified · 421 papers

Approaches in this cluster:

Most cited and most cited since 2024:

computer vision · humans · abstract · 365 papers

Approaches in this cluster:

Most cited and most cited since 2024:

concepts · brain · fmri · 210 papers

Approaches in this cluster:

Most cited and most cited since 2024:

grounding · referring · expression · 123 papers

Approaches in this cluster:

Most cited and most cited since 2024:

Related topics in Vision-language and multimodal