Research map / Vision-language and multimodal
Human and embodied modeling: research map
1,673 accepted papers on Human and embodied modeling in Vision-language and multimodal, from ICML, NeurIPS, ICLR, CVPR and AAAI (2016–2026), grouped into 5 clusters and 21 approaches. The busiest year so far is 2025.
Within Vision-language and multimodal, its share shrank from 33.5% in 2023–24 to 14.0% in 2025–26 (532 → 669 papers at ICML, NeurIPS, CVPR and AAAI, the venues with data for all four years).
Explore Human and embodied modeling in the interactive map
Working on something in this topic? Describe your idea in scime atlas to see which approach it falls under, the closest papers by meaning and how crowded the spot has become.
Approaches and key papers
retrieval · scene · captioning · 554 papers
Approaches in this cluster:
- Scene graph generation (126 papers)
Generates open-vocabulary and training-free scene graphs using vision-language models and object-centric binding. - Language-grounded scene text and structure (124 papers)
Connects images with text through scene text recognition, language bias correction, and scene-language representations. - Multimodal retrieval-augmented modeling (85 papers)
Uses cross-modal alignment, retrieval, and relation propagation for multimodal dialog and language modeling. - Composed and text-based image retrieval (78 papers)
Retrieves images from text, sketch, or reference-plus-modification queries with vision-language encoders. - Concept-guided image captioning (69 papers)
Improves captions by injecting semantic concepts, context guidance, and scene-graph evaluation. - Neural image captioning (72 papers)
Generate descriptive, controllable, and personalized captions from images.
Most cited and most cited since 2024:
- Image Captioning With Semantic Attention (CVPR 2016 · 1,739 citations)
- Knowing When to Look: Adaptive Attention via a Visual Sentinel for Image Captioning (CVPR 2017 · 1,734 citations)
- HIPTrack: Visual Tracking with Historical Prompts (CVPR 2024 · 105 citations)
- UFineBench: Towards Text-based Person Retrieval with Ultra-fine Granularity (CVPR 2024 · 68 citations)
pre · generation · unified · 421 papers
Approaches in this cluster:
- Vision-language pretraining (182 papers)
Pretrains unified vision-language models using image-as-language tokens and decoupled visual encoding. - Medical vision-language models (130 papers)
Pretrains medical vision-language models for radiology reports and 3D imaging. - Layout-aware document understanding (73 papers)
Extracts information from visually rich documents with layout-aware pretrained models. - Text-driven human motion generation (36 papers)
Generates human motion and interactions from text using language-aligned motion models.
Most cited and most cited since 2024:
- Deep contextualized word representations (ICLR 2018 · 12,263 citations)
- BEiT: BERT Pre-Training of Image Transformers (ICLR 2022 · 915 citations)
- SkySense: A Multi-Modal Remote Sensing Foundation Model Towards Universal Interpretation for Earth Observation Imagery (CVPR 2024 · 251 citations)
- Modeling Dense Multimodal Interactions Between Biological Pathways and Histology for Survival Prediction (CVPR 2024 · 143 citations)
computer vision · humans · abstract · 365 papers
Approaches in this cluster:
- Multimodal and affective interaction modeling (87 papers)
Models multimodal interactions, emotions, and social cues with datasets and information decomposition. - Human-like visual cognition in VLMs (97 papers)
Tests and improves vision-language models on human visual cognition, compositionality, and structure extraction. - Human perception modeling (71 papers)
Models human visual motion, saliency, and eye movement with trainable networks and human signals. - Abstract visual reasoning benchmarks (63 papers)
Benchmarks few-shot and systematic visual reasoning using object-centric abstraction and analogical learning. - Scientific domain vision datasets (47 papers)
Builds large-scale datasets for ecological, geological, and agricultural imagery plus physical prediction.
Most cited and most cited since 2024:
- VMamba: Visual State Space Model (NeurIPS 2024 · 830 citations)
- Detecting Visual Relationships With Deep Relational Networks (CVPR 2017 · 507 citations)
- BioCLIP: A Vision Foundation Model for the Tree of Life (CVPR 2024 · 124 citations)
- EmbodiedScan: A Holistic Multi-Modal 3D Perception Suite Towards Embodied AI (CVPR 2024 · 78 citations)
concepts · brain · fmri · 210 papers
Approaches in this cluster:
- fMRI visual decoding (106 papers)
Decodes and synthesizes visual stimuli from brain recordings using semantic and cross-subject models. - Language-grounded visual concept learning (83 papers)
Learns visual concepts by tokenization, decomposition, and transfer from language models. - Concept bottleneck models (21 papers)
Builds interpretable classifiers that predict human-understandable concepts using vision-language guidance.
Most cited and most cited since 2024:
- Dense Captioning With Joint Inference and Visual Context (CVPR 2017 · 175 citations)
- PaStaNet: Toward Human Activity Knowledge Engine (CVPR 2020 · 145 citations)
- MICA: Towards Explainable Skin Lesion Diagnosis via Multi-Level Image-Concept Alignment (AAAI 2024 · 27 citations)
- VLG-CBM: Training Concept Bottleneck Models with Vision-Language Guidance (NeurIPS 2024 · 21 citations)
grounding · referring · expression · 123 papers
Approaches in this cluster:
- Visual grounding transformers (63 papers)
Localizes language queries in images and 3D scenes with decoupled fusion and verification. - Region-level vision-language alignment (32 papers)
Aligns regions and objects with language through visual experts and hierarchical grounded annotation. - Referring expression comprehension (28 papers)
Resolves referring expressions with one-step transformers, iterative scanning, and adaptive inference.
Most cited and most cited since 2024:
- MAttNet: Modular Attention Network for Referring Expression Comprehension (CVPR 2018 · 869 citations)
- Modeling Relationships in Referential Expressions With Compositional Modular Networks (CVPR 2017 · 410 citations)
- GLaMM: Pixel Grounding Large Multimodal Model (CVPR 2024 · 175 citations)
- CogVLM: Visual Expert for Pretrained Language Models (NeurIPS 2024 · 77 citations)
Related topics in Vision-language and multimodal
- Visual question answering (546)
- Multimodal LLMs (2,371)
- Robot action and navigation (630)
- CLIP and zero-shot (959)
- Cross-modal learning (1,303)
