Research map / Vision-language and multimodal
Cross-modal learning: research map
1,303 accepted papers on Cross-modal learning in Vision-language and multimodal, from ICML, NeurIPS, ICLR, CVPR and AAAI (2016–2026), grouped into 3 clusters and 12 approaches. The busiest year so far is 2026.
Within Vision-language and multimodal, its share shrank from 18.3% in 2023–24 to 16.5% in 2025–26 (290 → 788 papers at ICML, NeurIPS, CVPR and AAAI, the venues with data for all four years).
Explore Cross-modal learning in the interactive map
Working on something in this topic? Describe your idea in scime atlas to see which approach it falls under, the closest papers by meaning and how crowded the spot has become.
Approaches and key papers
fusion · missing · multi modal · 714 papers
Approaches in this cluster:
- Multimodal fusion with missing modalities (208 papers)
Learns aligned multimodal representations robust to missing or imbalanced modalities. - Multi-modality image fusion (194 papers)
Fuses RGB, infrared, event and other modalities with transformers and semantic priors. - Medical multimodal fusion (166 papers)
Combines clinical and imaging modalities with decomposition, symmetric consistency and graph structure. - Multimodal sentiment analysis (106 papers)
Uses distillation and consistency learning for sentiment recognition, including with missing modalities. - Multi-modal object re-identification (40 papers)
Selects tokens and decouples features across RGB, infrared and text for re-identification.
Most cited and most cited since 2024:
- Learning Modality-Specific Representations with Self-Supervised Multi-Task Learning for Multimodal Sentiment Analysis (AAAI 2021 · 751 citations)
- What Makes Training Multi-Modal Classification Networks Hard? (CVPR 2020 · 532 citations)
- Equivariant Multi-Modality Image Fusion (CVPR 2024 · 233 citations)
- Bi-directional Adapter for Multimodal Tracking (AAAI 2024 · 132 citations)
cross · retrieval · image text · 402 papers
Approaches in this cluster:
- Unimodal-to-multimodal representation transfer (145 papers)
Studies cross-modal misalignment, distillation and dataset condensation to learn across modalities from unimodal data. - Probabilistic embeddings for cross-modal retrieval (108 papers)
Matches images and text using probabilistic, multi-embedding or generative-aided representations. - Contrastive vision-language alignment (104 papers)
Refines contrastive and preference-based alignment between image and text embeddings. - Cross-modal hashing (45 papers)
Learns compact binary codes for efficient cross-modal retrieval.
Most cited and most cited since 2024:
- Deep Cross-Modal Hashing (CVPR 2017 · 898 citations)
- ImageBind: One Embedding Space To Bind Them All (CVPR 2023 · 810 citations)
- Noisy-Correspondence Learning for Text-to-Image Person Re-identification (CVPR 2024 · 126 citations)
- CricaVPR: Cross-image Correlation-aware Representation Learning for Visual Place Recognition (CVPR 2024 · 78 citations)
audio · sound · speech · 187 papers
Approaches in this cluster:
- Audio-visual source localization (67 papers)
Aligns sound and visual scenes to localize and segment sounding objects. - Audio-visual speech separation (76 papers)
Uses visual cues such as lips and cross-modal prompts to separate speech and adapt pretrained models. - Self-supervised audio-visual representation learning (44 papers)
Learns audio-visual features from correspondence and synchronization via contrastive and distillation objectives.
Most cited and most cited since 2024:
- Balanced Multimodal Learning via On-the-Fly Gradient Modulation (CVPR 2022 · 336 citations)
- Seeing Voices and Hearing Faces: Cross-Modal Biometric Matching (CVPR 2018 · 240 citations)
- AVFF: Audio-Visual Feature Fusion for Video Deepfake Detection (CVPR 2024 · 76 citations)
- AVSegFormer: Audio-Visual Segmentation with Transformer (AAAI 2024 · 67 citations)
Related topics in Vision-language and multimodal
- Human and embodied modeling (1,673)
- Visual question answering (546)
- Multimodal LLMs (2,371)
- Robot action and navigation (630)
- CLIP and zero-shot (959)
