Research map / Adversarial robustness and security
Attacks on text and graphs: research map
1,088 accepted papers on Attacks on text and graphs in Adversarial robustness and security, from ICML, NeurIPS, ICLR, CVPR and AAAI (2016–2026), grouped into 4 clusters and 14 approaches. The busiest year so far is 2026.
Within Adversarial robustness and security, its share grew from 35.9% in 2023–24 to 49.5% in 2025–26 (279 → 486 papers at ICML, NeurIPS, CVPR and AAAI, the venues with data for all four years).
Explore Attacks on text and graphs in the interactive map
Working on something in this topic? Describe your idea in scime atlas to see which approach it falls under, the closest papers by meaning and how crowded the spot has become.
Approaches and key papers
generative · purification · diffusion · 342 papers
Approaches in this cluster:
- Defenses against universal perturbations (97 papers)
Detect, denoise or generative-purify adversarial examples and perturbations. - Attacks and defenses in sequential settings (96 papers)
Study adversarial attacks on RL, online learners and foundation models. - Diffusion-based adversarial purification (78 papers)
Purify or generate adversarial examples using diffusion models. - Graph neural network robustness (71 papers)
Attack and defend graph neural networks against adversarial perturbations.
Most cited and most cited since 2024:
- Universal Adversarial Perturbations (CVPR 2017 · 2,731 citations)
- Multi-Adversarial Discriminative Deep Domain Generalization for Face Presentation Attack Detection (CVPR 2019 · 420 citations)
- Adv-Diffusion: Imperceptible Adversarial Face Identity Attack via Latent Diffusion Model (AAAI 2024 · 35 citations)
- Pre-trained Model Guided Fine-Tuning for Zero-Shot Adversarial Robustness (CVPR 2024 · 27 citations)
llms · prompt · jailbreak · 307 papers
Approaches in this cluster:
- Adversarial prompts for language models (98 papers)
Attack and defend language models with adversarial prompts and robustness benchmarks. - LLM jailbreak and injection defense (113 papers)
Study jailbreaks and prompt injections and train or evaluate safer LLMs. - Vision-language model robustness (96 papers)
Attack and harden vision-language models with multimodal adversarial search and prompt tuning.
Most cited and most cited since 2024:
- FreeLB: Enhanced Adversarial Training for Natural Language Understanding (ICLR 2020 · 286 citations)
- Visual Adversarial Examples Jailbreak Aligned Large Language Models (AAAI 2024 · 122 citations)
- AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents (NeurIPS 2024 · 83 citations)
- JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models (NeurIPS 2024 · 78 citations)
physical · object · patches · 227 papers
Approaches in this cluster:
- Biologically inspired robust vision (77 papers)
Improve adversarial robustness with foveation, visual-cortex front-ends, nearest-neighbor defenses and unadversarial objects. - Adversarial perturbations beyond images (69 papers)
Craft universal, video and physical-world adversarial examples and self-supervised defenses. - Physical adversarial patches and camouflage (56 papers)
Design and defend against printable patches and 3D camouflage fooling detectors. - Latency and driving-perception attacks (25 papers)
Attack autonomous-driving perception by increasing latency or fooling multiple modules.
Most cited and most cited since 2024:
- Robust Physical-World Attacks on Deep Learning Visual Classification (CVPR 2018 · 2,168 citations)
- Synthesizing Robust Adversarial Examples (ICML 2018 · 667 citations)
- Physical 3D Adversarial Attacks against Monocular Depth Estimation in Autonomous Driving (CVPR 2024 · 54 citations)
- PAD: Patch-Agnostic Defense against Adversarial Patch Attacks (CVPR 2024 · 44 citations)
poisoning · membership · watermarking · 212 papers
Approaches in this cluster:
- Membership inference and inversion (118 papers)
Attack and defend models through membership inference, model inversion and reconstruction. - Neural network watermarking (53 papers)
Embed watermarks and ownership proofs in models to protect intellectual property. - Clean-label data poisoning (41 papers)
Craft poisoned training data via gradient matching and generative methods.
Most cited and most cited since 2024:
- The Secret Revealer: Generative Model-Inversion Attacks Against Deep Neural Networks (CVPR 2020 · 442 citations)
- Analyzing Federated Learning through an Adversarial Lens (ICML 2019 · 382 citations)
- Learning to Unlearn: Instance-Wise Unlearning for Pre-trained Classifiers (AAAI 2024 · 40 citations)
- Watermark-embedded Adversarial Examples for Copyright Protection against Diffusion Models (CVPR 2024 · 20 citations)
