Research map / Vision-language and multimodal
CLIP and zero-shot: research map
959 accepted papers on CLIP and zero-shot in Vision-language and multimodal, from ICML, NeurIPS, ICLR, CVPR and AAAI (2016–2026), grouped into 3 clusters and 12 approaches. The busiest year so far is 2026.
Within Vision-language and multimodal, its share shrank from 23.4% in 2023–24 to 10.5% in 2025–26 (371 → 500 papers at ICML, NeurIPS, CVPR and AAAI, the venues with data for all four years).
Explore CLIP and zero-shot in the interactive map
Working on something in this topic? Describe your idea in scime atlas to see which approach it falls under, the closest papers by meaning and how crowded the spot has become.
Approaches and key papers
zero shot · open · vocabulary · 417 papers
Approaches in this cluster:
- Open-vocabulary alignment from VLMs (130 papers)
Refines vision-language alignment and similarity scoring to enable open-vocabulary recognition and segmentation. - Attribute-based zero-shot learning (115 papers)
Uses attribute-guided transformers, prototypes and semantic grounding to classify unseen classes. - Adapting VLMs for zero-shot classification (108 papers)
Improves zero-shot classification by adding text descriptions, LLM prompts and long-tail handling to VLMs. - Language models for zero-shot reasoning (64 papers)
Leverages pretrained language models as priors or components for zero-shot multimodal and cross-lingual tasks.
Most cited and most cited since 2024:
- Learning Transferable Visual Models From Natural Language Supervision (ICML 2021 · 5,285 citations)
- Matching Networks for One Shot Learning (NeurIPS 2016 · 923 citations)
- AnomalyGPT: Detecting Industrial Anomalies Using Large Vision-Language Models (AAAI 2024 · 279 citations)
- Florence-2: Advancing a Unified Representation for a Variety of Vision Tasks (CVPR 2024 · 264 citations)
prompt · tuning · pre trained · 286 papers
Approaches in this cluster:
- Prompt learning for CLIP (102 papers)
Learns soft text prompts for vision-language models, with distillation, regularization and Bayesian modeling. - Test-time adaptation of VLMs (93 papers)
Adapts vision-language models on unlabeled test data via prompt tuning and lightweight adapters. - Feature calibration for unseen-class detection (50 papers)
Calibrates multimodal features to handle few-shot, unseen-class and out-of-distribution inputs. - Parameter-efficient visual tuning (41 papers)
Adapts pretrained vision and multimodal models through visual prompt tuning or efficient fine-tuning.
Most cited and most cited since 2024:
- Conditional Prompt Learning for Vision-Language Models (CVPR 2022 · 1,735 citations)
- MaPLe: Multi-Modal Prompt Learning (CVPR 2023 · 886 citations)
- PromptKD: Unsupervised Prompt Distillation for Vision-Language Models (CVPR 2024 · 106 citations)
- MmAP: Multi-Modal Alignment Prompt for Cross-Domain Multi-Task Learning (AAAI 2024 · 79 citations)
contrastive · language image · image text · 256 papers
Approaches in this cluster:
- CLIP training recipes and data curation (72 papers)
Improves contrastive language-image pretraining through data scaling, curation and added supervision. - Recaptioning and text-supervised pretraining (73 papers)
Uses synthetic captions and interleaved image-text data to learn better visual representations. - Fine-grained improvements to CLIP (63 papers)
Strengthens CLIP's local detail understanding and studies its scaling behaviour. - Self-distillation for vision-language pretraining (48 papers)
Pretrains vision-language models with bootstrapping, masked and self-distillation objectives.
Most cited and most cited since 2024:
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision (ICML 2021 · 1,195 citations)
- LAION-5B: An open large-scale dataset for training next generation image-text models (NeurIPS 2022 · 1,032 citations)
- VadCLIP: Adapting Vision-Language Models for Weakly Supervised Video Anomaly Detection (AAAI 2024 · 206 citations)
- An Empirical Study of CLIP for Text-Based Person Search (AAAI 2024 · 99 citations)
Related topics in Vision-language and multimodal
- Human and embodied modeling (1,673)
- Visual question answering (546)
- Multimodal LLMs (2,371)
- Robot action and navigation (630)
- Cross-modal learning (1,303)
