atlas

Research map / Vision-language and multimodal

CLIP and zero-shot: research map

959 accepted papers on CLIP and zero-shot in Vision-language and multimodal, from ICML, NeurIPS, ICLR, CVPR and AAAI (2016–2026), grouped into 3 clusters and 12 approaches. The busiest year so far is 2026.

Within Vision-language and multimodal, its share shrank from 23.4% in 2023–24 to 10.5% in 2025–26 (371 → 500 papers at ICML, NeurIPS, CVPR and AAAI, the venues with data for all four years).

2016: 5162017: 2172018: 5182019: 3192020: 6202021: 8212022: 59222023: 125232024: 246242025: 245252026: 25526

Explore CLIP and zero-shot in the interactive map

Working on something in this topic? Describe your idea in scime atlas to see which approach it falls under, the closest papers by meaning and how crowded the spot has become.

Approaches and key papers

zero shot · open · vocabulary · 417 papers

Approaches in this cluster:

Most cited and most cited since 2024:

prompt · tuning · pre trained · 286 papers

Approaches in this cluster:

Most cited and most cited since 2024:

contrastive · language image · image text · 256 papers

Approaches in this cluster:

Most cited and most cited since 2024:

Related topics in Vision-language and multimodal