Research map / Vision-language and multimodal
Multimodal LLMs: research map
2,371 accepted papers on Multimodal LLMs in Vision-language and multimodal, from ICML, NeurIPS, ICLR, CVPR and AAAI (2016–2026), grouped into 4 clusters and 15 approaches. The busiest year so far is 2026.
Within Vision-language and multimodal, its share grew from 13.0% in 2023–24 to 45.1% in 2025–26 (207 → 2,156 papers at ICML, NeurIPS, CVPR and AAAI, the venues with data for all four years).
Explore Multimodal LLMs in the interactive map
Working on something in this topic? Describe your idea in scime atlas to see which approach it falls under, the closest papers by meaning and how crowded the spot has become.
Approaches and key papers
vlms · visual reasoning · chain · 1,218 papers
Approaches in this cluster:
- VLM capability analysis (264 papers)
Probes and refines vision-language models, assessing priors, reasoning, and design choices. - Chain-of-thought multimodal reasoning (332 papers)
Improves multimodal LLM reasoning with latent thinking, long chains, and structured scene-graph reasoning. - Multimodal LLM benchmarks (252 papers)
Builds comprehensive and evolving benchmarks for large multimodal models. - Spatial reasoning in VLMs (207 papers)
Improves grounded spatial reasoning and perspective understanding in vision-language models. - Reinforcement learning for visual reasoning (163 papers)
Trains vision-language models with reinforcement learning and data augmentation to reason visually.
Most cited and most cited since 2024:
- MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI (CVPR 2024 · 355 citations)
- SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities (CVPR 2024 · 234 citations)
- SpatialRGPT: Grounded Spatial Reasoning in Vision-Language Models (NeurIPS 2024 · 98 citations)
- Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces (CVPR 2025 · 92 citations)
mllm · instruction · llms · 753 papers
Approaches in this cluster:
- Instruction-tuned multimodal assistants (199 papers)
Builds stronger MLLMs through better visual encoders, small language models, and fine-grained visual knowledge or preference training. - Benchmarks and analysis of MLLMs (181 papers)
Evaluates and probes multimodal LLMs, covering feature fusion, information flow, benchmarks, and jailbreak weaknesses. - Unified any-modality language models (152 papers)
Aligns audio, video and other modalities with LLMs in one framework while avoiding text-only forgetting. - Interleaved image-text generation with MLLMs (153 papers)
Uses multimodal LLMs to jointly understand and generate images and text in context. - Visual instruction tuning (68 papers)
Fine-tunes language models on image-instruction data, generated or filtered, to follow multimodal instructions.
Most cited and most cited since 2024:
- Improved Baselines with Visual Instruction Tuning (CVPR 2024 · 1,490 citations)
- Visual Instruction Tuning (NeurIPS 2023 · 1,198 citations)
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models (ICLR 2024 · 480 citations)
- CogAgent: A Visual Language Model for GUI Agents (CVPR 2024 · 173 citations)
tokens · pruning · visual token · 207 papers
Approaches in this cluster:
- Attention-based visual token pruning (117 papers)
Drops redundant visual tokens inside VLMs using attention or text-guided importance scores to speed up inference. - Training-free token reduction for MLLMs (90 papers)
Searches or selects which vision tokens to trim or merge across layers and frames without retraining.
Most cited and most cited since 2024:
- VisionZip: Longer is Better but Not Necessary in Vision Language Models (CVPR 2025 · 46 citations)
- NVILA: Efficient Frontier Visual Language Models (CVPR 2025 · 37 citations)
- FastVLM: Efficient Vision Encoding for Vision Language Models (CVPR 2025 · 27 citations)
- DivPrune: Diversity-based Visual Token Pruning for Large Multimodal Models (CVPR 2025 · 27 citations)
hallucination · lvlms · mitigating · 193 papers
Approaches in this cluster:
- Decoding-time hallucination mitigation (91 papers)
Alters decoding, such as contrastive or layer-wise decoding, to reduce object hallucination in vision-language models. - Preference tuning against hallucination (58 papers)
Trains models with preference optimization and targeted data to suppress hallucinated content and locate its causes. - Bias-aware multi-image robustness (44 papers)
Analyzes language and position biases in LVLMs and corrects them through selective training or attention perturbation.
Most cited and most cited since 2024:
- Mitigating Object Hallucinations in Large Vision-Language Models through Visual Contrastive Decoding (CVPR 2024 · 172 citations)
- Detecting and Preventing Hallucinations in Large Vision Language Models (AAAI 2024 · 145 citations)
- HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models (CVPR 2024 · 138 citations)
- OPERA: Alleviating Hallucination in Multi-Modal Large Language Models via Over-Trust Penalty and Retrospection-Allocation (CVPR 2024 · 124 citations)
Related topics in Vision-language and multimodal
- Human and embodied modeling (1,673)
- Visual question answering (546)
- Robot action and navigation (630)
- CLIP and zero-shot (959)
- Cross-modal learning (1,303)
