A Systematic Review of Poisoning Attacks Against Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Fendley, Neil, Staley, Edward W., Carney, Joshua, Redman, William, Chau, Marie, Drenkow, Nathan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Backdoors in DRL: Four Environments Focusing on In-distribution Triggers
by: Ashcraft, Chace, et al.
Published: (2025)
by: Ashcraft, Chace, et al.
Published: (2025)
Poison with Style: A Practical Poisoning Attack on Code Large Language Models
by: Tran, Khang, et al.
Published: (2026)
by: Tran, Khang, et al.
Published: (2026)
How to Defend Against Large-scale Model Poisoning Attacks in Federated Learning: A Vertical Solution
by: Wang, Jinbo, et al.
Published: (2024)
by: Wang, Jinbo, et al.
Published: (2024)
PoisonedParrot: Subtle Data Poisoning Attacks to Elicit Copyright-Infringing Content from Large Language Models
by: Panaitescu-Liess, Michael-Andrei, et al.
Published: (2025)
by: Panaitescu-Liess, Michael-Andrei, et al.
Published: (2025)
Shadowcast: Stealthy Data Poisoning Attacks Against Vision-Language Models
by: Xu, Yuancheng, et al.
Published: (2024)
by: Xu, Yuancheng, et al.
Published: (2024)
Poisoned-MRAG: Knowledge Poisoning Attacks to Multimodal Retrieval Augmented Generation
by: Liu, Yinuo, et al.
Published: (2025)
by: Liu, Yinuo, et al.
Published: (2025)
PoisonedRAG: Knowledge Corruption Attacks to Retrieval-Augmented Generation of Large Language Models
by: Zou, Wei, et al.
Published: (2024)
by: Zou, Wei, et al.
Published: (2024)
SecureLearn -- An Attack-agnostic Defense for Multiclass Machine Learning Against Data Poisoning Attacks
by: Paracha, Anum, et al.
Published: (2025)
by: Paracha, Anum, et al.
Published: (2025)
SleeperNets: Universal Backdoor Poisoning Attacks Against Reinforcement Learning Agents
by: Rathbun, Ethan, et al.
Published: (2024)
by: Rathbun, Ethan, et al.
Published: (2024)
Provable Robustness of (Graph) Neural Networks Against Data Poisoning and Backdoor Attacks
by: Gosch, Lukas, et al.
Published: (2024)
by: Gosch, Lukas, et al.
Published: (2024)
Defending Against Sophisticated Poisoning Attacks with RL-based Aggregation in Federated Learning
by: Wang, Yujing, et al.
Published: (2024)
by: Wang, Yujing, et al.
Published: (2024)
Online Poisoning Attack Against Reinforcement Learning under Black-box Environments
by: Li, Jianhui, et al.
Published: (2024)
by: Li, Jianhui, et al.
Published: (2024)
FedRecAttack: Model Poisoning Attack to Federated Recommendation
by: Rong, Dazhong, et al.
Published: (2022)
by: Rong, Dazhong, et al.
Published: (2022)
Systematic Scaling Analysis of Jailbreak Attacks in Large Language Models
by: Wang, Xiangwen, et al.
Published: (2026)
by: Wang, Xiangwen, et al.
Published: (2026)
Defending Large Language Models Against Attacks With Residual Stream Activation Analysis
by: Kawasaki, Amelia, et al.
Published: (2024)
by: Kawasaki, Amelia, et al.
Published: (2024)
Ensembling Membership Inference Attacks Against Tabular Generative Models
by: Ward, Joshua, et al.
Published: (2025)
by: Ward, Joshua, et al.
Published: (2025)
Transferable Availability Poisoning Attacks
by: Liu, Yiyong, et al.
Published: (2023)
by: Liu, Yiyong, et al.
Published: (2023)
Concept-ROT: Poisoning Concepts in Large Language Models with Model Editing
by: Grimes, Keltin, et al.
Published: (2024)
by: Grimes, Keltin, et al.
Published: (2024)
Composite Backdoor Attacks Against Large Language Models
by: Huang, Hai, et al.
Published: (2023)
by: Huang, Hai, et al.
Published: (2023)
Provable Watermarking for Data Poisoning Attacks
by: Zhu, Yifan, et al.
Published: (2025)
by: Zhu, Yifan, et al.
Published: (2025)
Poisoning the Watchtower: Prompt Injection Attacks Against LLM-Augmented Security Operations Through Adversarial Log Content
by: Pandey, Rohan, et al.
Published: (2026)
by: Pandey, Rohan, et al.
Published: (2026)
GShield: Mitigating Poisoning Attacks in Federated Learning
by: M., Sameera K., et al.
Published: (2025)
by: M., Sameera K., et al.
Published: (2025)
Indiscriminate Data Poisoning Attacks on Neural Networks
by: Lu, Yiwei, et al.
Published: (2022)
by: Lu, Yiwei, et al.
Published: (2022)
Adaptive and Robust Data Poisoning Detection and Sanitization in Wearable IoT Systems using Large Language Models
by: Mithsara, W. K. M, et al.
Published: (2025)
by: Mithsara, W. K. M, et al.
Published: (2025)
Attacking LLMs and AI Agents: Advertisement Embedding Attacks Against Large Language Models
by: Guo, Qiming, et al.
Published: (2025)
by: Guo, Qiming, et al.
Published: (2025)
Temporal Context Awareness: A Defense Framework Against Multi-turn Manipulation Attacks on Large Language Models
by: Kulkarni, Prashant, et al.
Published: (2025)
by: Kulkarni, Prashant, et al.
Published: (2025)
State Backdoor: Towards Stealthy Real-world Poisoning Attack on Vision-Language-Action Model in State Space
by: Guo, Ji, et al.
Published: (2026)
by: Guo, Ji, et al.
Published: (2026)
Link Stealing Attacks Against Inductive Graph Neural Networks
by: Wu, Yixin, et al.
Published: (2024)
by: Wu, Yixin, et al.
Published: (2024)
Comments on "Privacy-Enhanced Federated Learning Against Poisoning Adversaries"
by: Schneider, Thomas, et al.
Published: (2024)
by: Schneider, Thomas, et al.
Published: (2024)
Exact Certification of (Graph) Neural Networks Against Label Poisoning
by: Sabanayagam, Mahalakshmi, et al.
Published: (2024)
by: Sabanayagam, Mahalakshmi, et al.
Published: (2024)
Thought-Transfer: Indirect Targeted Poisoning Attacks on Chain-of-Thought Reasoning Models
by: Chaudhari, Harsh, et al.
Published: (2026)
by: Chaudhari, Harsh, et al.
Published: (2026)
Local Environment Poisoning Attacks on Federated Reinforcement Learning
by: Ma, Evelyn, et al.
Published: (2023)
by: Ma, Evelyn, et al.
Published: (2023)
Inverting Gradient Attacks Makes Powerful Data Poisoning
by: Bouaziz, Wassim, et al.
Published: (2024)
by: Bouaziz, Wassim, et al.
Published: (2024)
FedMID: A Data-Free Method for Using Intermediate Outputs as a Defense Mechanism Against Poisoning Attacks in Federated Learning
by: Han, Sungwon, et al.
Published: (2024)
by: Han, Sungwon, et al.
Published: (2024)
Mirror Mirror on the Wall, Have I Forgotten it All? A New Framework for Evaluating Machine Unlearning
by: Brimhall, Brennon, et al.
Published: (2025)
by: Brimhall, Brennon, et al.
Published: (2025)
Defending Against Poisoning Attacks in Federated Learning with Blockchain
by: Dong, Nanqing, et al.
Published: (2023)
by: Dong, Nanqing, et al.
Published: (2023)
Indiscriminate Data Poisoning Attacks on Pre-trained Feature Extractors
by: Lu, Yiwei, et al.
Published: (2024)
by: Lu, Yiwei, et al.
Published: (2024)
Data Poisoning Attacks in Intelligent Transportation Systems: A Survey
by: Wang, Feilong, et al.
Published: (2024)
by: Wang, Feilong, et al.
Published: (2024)
Sybil-based Virtual Data Poisoning Attacks in Federated Learning
by: Zhu, Changxun, et al.
Published: (2025)
by: Zhu, Changxun, et al.
Published: (2025)
Debiased Graph Poisoning Attack via Contrastive Surrogate Objective
by: Yoon, Kanghoon, et al.
Published: (2024)
by: Yoon, Kanghoon, et al.
Published: (2024)
Similar Items
-
Backdoors in DRL: Four Environments Focusing on In-distribution Triggers
by: Ashcraft, Chace, et al.
Published: (2025) -
Poison with Style: A Practical Poisoning Attack on Code Large Language Models
by: Tran, Khang, et al.
Published: (2026) -
How to Defend Against Large-scale Model Poisoning Attacks in Federated Learning: A Vertical Solution
by: Wang, Jinbo, et al.
Published: (2024) -
PoisonedParrot: Subtle Data Poisoning Attacks to Elicit Copyright-Infringing Content from Large Language Models
by: Panaitescu-Liess, Michael-Andrei, et al.
Published: (2025) -
Shadowcast: Stealthy Data Poisoning Attacks Against Vision-Language Models
by: Xu, Yuancheng, et al.
Published: (2024)