Research map / Vision-language and multimodal
Robot action and navigation: research map
630 accepted papers on Robot action and navigation in Vision-language and multimodal, from ICML, NeurIPS, ICLR, CVPR and AAAI (2016–2026), grouped into 2 clusters and 6 approaches. The busiest year so far is 2026.
Within Vision-language and multimodal, its share grew from 4.8% in 2023–24 to 10.4% in 2025–26 (76 → 497 papers at ICML, NeurIPS, CVPR and AAAI, the venues with data for all four years).
Explore Robot action and navigation in the interactive map
Working on something in this topic? Describe your idea in scime atlas to see which approach it falls under, the closest papers by meaning and how crowded the spot has become.
Approaches and key papers
vla · language action · robotic · 487 papers
Approaches in this cluster:
- Vision-language-action foundation models (264 papers)
Trains and benchmarks end-to-end VLA policies, with reasoning, unified representations and efficient token caching. - Memory and subgoal-based manipulation (103 papers)
Improves robot manipulation with memory, behavioral phasing, subgoal images and skill decomposition. - Embodied multimodal LLM agents (70 papers)
Applies pretrained multimodal LLMs and embodied chain-of-thought to interactive agents, with benchmarks. - Multimodal prompt-driven manipulation (50 papers)
Conditions robot policies on multimodal prompts and object-centric or programmatic plans from language instructions.
Most cited and most cited since 2024:
- ALFRED: A Benchmark for Interpreting Grounded Instructions for Everyday Tasks (CVPR 2020 · 533 citations)
- PaLM-E: An Embodied Multimodal Language Model (ICML 2023 · 349 citations)
- ManipLLM: Embodied Multimodal Large Language Model for Object-Centric Robotic Manipulation (CVPR 2024 · 82 citations)
- Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision Language Audio and Action (CVPR 2024 · 77 citations)
language navigation · vln · agent · 143 papers
Approaches in this cluster:
- Pretraining and data augmentation for VLN (97 papers)
Improves navigation agents through pretraining, synthetic instructions, imitation learning and exploration with scene objects. - Reasoning-enhanced vision-language navigation (46 papers)
Adds auxiliary tasks, knowledge, concept learning and structured state modeling to navigation policies.
Most cited and most cited since 2024:
- Vision-and-Language Navigation: Interpreting Visually-Grounded Navigation Instructions in Real Environments (CVPR 2018 · 1,195 citations)
- Reinforced Cross-Modal Matching and Self-Supervised Imitation Learning for Vision-Language Navigation (CVPR 2019 · 548 citations)
- NavGPT: Explicit Reasoning in Vision-and-Language Navigation with Large Language Models (AAAI 2024 · 184 citations)
- Vision-and-Language Navigation via Causal Learning (CVPR 2024 · 40 citations)
Related topics in Vision-language and multimodal
- Human and embodied modeling (1,673)
- Visual question answering (546)
- Multimodal LLMs (2,371)
- CLIP and zero-shot (959)
- Cross-modal learning (1,303)
