Scaling Trends for Data Poisoning in LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Bowen, Dillon, Murphy, Brendan, Cai, Will, Khachaturov, David, Gleave, Adam, Pelrine, Kellin |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Exposing the Systematic Vulnerability of Open-Weight Models to Prefill Attacks
by: Struppek, Lukas, et al.
Published: (2026)
by: Struppek, Lukas, et al.
Published: (2026)
Jailbreak-Tuning: Models Efficiently Learn Jailbreak Susceptibility
by: Murphy, Brendan, et al.
Published: (2025)
by: Murphy, Brendan, et al.
Published: (2025)
Exploiting Novel GPT-4 APIs
by: Pelrine, Kellin, et al.
Published: (2023)
by: Pelrine, Kellin, et al.
Published: (2023)
Complexity Matters: Effective Dimensionality as a Measure for Adversarial Robustness
by: Khachaturov, David, et al.
Published: (2024)
by: Khachaturov, David, et al.
Published: (2024)
Have You Poisoned My Data? Defending Neural Networks against Data Poisoning
by: De Gaspari, Fabio, et al.
Published: (2024)
by: De Gaspari, Fabio, et al.
Published: (2024)
Scaling Trends in Language Model Robustness
by: Howe, Nikolaus, et al.
Published: (2024)
by: Howe, Nikolaus, et al.
Published: (2024)
Adversarial Suffix Filtering: a Defense Pipeline for LLMs
by: Khachaturov, David, et al.
Published: (2025)
by: Khachaturov, David, et al.
Published: (2025)
Data Poisoning Attacks on Off-Policy Policy Evaluation Methods
by: Lobo, Elita, et al.
Published: (2024)
by: Lobo, Elita, et al.
Published: (2024)
Silent Sabotage During Fine-Tuning: Few-Shot Rationale Poisoning of Compact Medical LLMs
by: Xie, Jingyuan, et al.
Published: (2026)
by: Xie, Jingyuan, et al.
Published: (2026)
Shadowcast: Stealthy Data Poisoning Attacks Against Vision-Language Models
by: Xu, Yuancheng, et al.
Published: (2024)
by: Xu, Yuancheng, et al.
Published: (2024)
Concealing Backdoor Model Updates in Federated Learning by Trigger-Optimized Data Poisoning
by: Zhang, Yujie, et al.
Published: (2024)
by: Zhang, Yujie, et al.
Published: (2024)
FedNIA: Noise-Induced Activation Analysis for Mitigating Data Poisoning in FL
by: Hallaji, Ehsan, et al.
Published: (2025)
by: Hallaji, Ehsan, et al.
Published: (2025)
SAFELOC: Overcoming Data Poisoning Attacks in Heterogeneous Federated Machine Learning for Indoor Localization
by: Singampalli, Akhil, et al.
Published: (2024)
by: Singampalli, Akhil, et al.
Published: (2024)
Be Kind, Rewrite: Benign Projections via Rewriting Defend Against LLM Data Poisoning Attacks
by: Halloran, John T., et al.
Published: (2026)
by: Halloran, John T., et al.
Published: (2026)
PureGen: Universal Data Purification for Train-Time Poison Defense via Generative Model Dynamics
by: Bhat, Sunay, et al.
Published: (2024)
by: Bhat, Sunay, et al.
Published: (2024)
GFCL: A GRU-based Federated Continual Learning Framework against Data Poisoning Attacks in IoV
by: Talpur, Anum, et al.
Published: (2022)
by: Talpur, Anum, et al.
Published: (2022)
Enhancing Prompt Injection Attacks to LLMs via Poisoning Alignment
by: Shao, Zedian, et al.
Published: (2024)
by: Shao, Zedian, et al.
Published: (2024)
Stealthy Poisoning Attacks Bypass Defenses in Regression Settings
by: Carnerero-Cano, Javier, et al.
Published: (2026)
by: Carnerero-Cano, Javier, et al.
Published: (2026)
SoK: Benchmarking Poisoning Attacks and Defenses in Federated Learning
by: Zhang, Heyi, et al.
Published: (2025)
by: Zhang, Heyi, et al.
Published: (2025)
KEPo: Knowledge Evolution Poison on Graph-based Retrieval-Augmented Generation
by: Chen, Qizhi, et al.
Published: (2026)
by: Chen, Qizhi, et al.
Published: (2026)
Semantic Chameleon: Corpus-Dependent Poisoning Attacks and Defenses in RAG Systems
by: Thornton, Scott
Published: (2026)
by: Thornton, Scott
Published: (2026)
Machine Unlearning Fails to Remove Data Poisoning Attacks
by: Pawelczyk, Martin, et al.
Published: (2024)
by: Pawelczyk, Martin, et al.
Published: (2024)
Universal Black-Box Reward Poisoning Attack against Offline Reinforcement Learning
by: Xu, Yinglun, et al.
Published: (2024)
by: Xu, Yinglun, et al.
Published: (2024)
ACE: A Model Poisoning Attack on Contribution Evaluation Methods in Federated Learning
by: Xu, Zhangchen, et al.
Published: (2024)
by: Xu, Zhangchen, et al.
Published: (2024)
FedReview: A Review Mechanism for Rejecting Poisoned Updates in Federated Learning
by: Zheng, Tianhang, et al.
Published: (2024)
by: Zheng, Tianhang, et al.
Published: (2024)
MemoryGraft: Persistent Compromise of LLM Agents via Poisoned Experience Retrieval
by: Srivastava, Saksham Sahai, et al.
Published: (2025)
by: Srivastava, Saksham Sahai, et al.
Published: (2025)
BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF
by: Duan, Kaiwen, et al.
Published: (2025)
by: Duan, Kaiwen, et al.
Published: (2025)
Reasoning Introduces New Poisoning Attacks Yet Makes Them More Complicated
by: Foerster, Hanna, et al.
Published: (2025)
by: Foerster, Hanna, et al.
Published: (2025)
Can Go AIs be adversarially robust?
by: Tseng, Tom, et al.
Published: (2024)
by: Tseng, Tom, et al.
Published: (2024)
BadSampler: Harnessing the Power of Catastrophic Forgetting to Poison Byzantine-robust Federated Learning
by: Liu, Yi, et al.
Published: (2024)
by: Liu, Yi, et al.
Published: (2024)
EAB-FL: Exacerbating Algorithmic Bias through Model Poisoning Attacks in Federated Learning
by: Meerza, Syed Irfan Ali, et al.
Published: (2024)
by: Meerza, Syed Irfan Ali, et al.
Published: (2024)
Best-of-Venom: Attacking RLHF by Injecting Poisoned Preference Data
by: Baumgärtner, Tim, et al.
Published: (2024)
by: Baumgärtner, Tim, et al.
Published: (2024)
Hidden Poison: Machine Unlearning Enables Camouflaged Poisoning Attacks
by: Di, Jimmy Z., et al.
Published: (2022)
by: Di, Jimmy Z., et al.
Published: (2022)
Tracing the Data Trail: A Survey of Data Provenance, Transparency and Traceability in LLMs
by: Hohensinner, Richard, et al.
Published: (2026)
by: Hohensinner, Richard, et al.
Published: (2026)
Fast Exact Unlearning for In-Context Learning Data for LLMs
by: Muresanu, Andrei I., et al.
Published: (2024)
by: Muresanu, Andrei I., et al.
Published: (2024)
Beyond Training-time Poisoning: Component-level and Post-training Backdoors in Deep Reinforcement Learning
by: Vyas, Sanyam, et al.
Published: (2025)
by: Vyas, Sanyam, et al.
Published: (2025)
OEP: Poisoning Self-Evolving LLM Agents via Locally Correct but Non-Transferable Experiences
by: Wang, Kaixiang, et al.
Published: (2026)
by: Wang, Kaixiang, et al.
Published: (2026)
Protecting Federated Learning from Extreme Model Poisoning Attacks via Multidimensional Time Series Anomaly Detection
by: Gabrielli, Edoardo, et al.
Published: (2023)
by: Gabrielli, Edoardo, et al.
Published: (2023)
Data-Free Privacy-Preserving for LLMs via Model Inversion and Selective Unlearning
by: Zhou, Xinjie, et al.
Published: (2026)
by: Zhou, Xinjie, et al.
Published: (2026)
Private-RAG: Answering Multiple Queries with LLMs while Keeping Your Data Private
by: Wu, Ruihan, et al.
Published: (2025)
by: Wu, Ruihan, et al.
Published: (2025)
Similar Items
-
Exposing the Systematic Vulnerability of Open-Weight Models to Prefill Attacks
by: Struppek, Lukas, et al.
Published: (2026) -
Jailbreak-Tuning: Models Efficiently Learn Jailbreak Susceptibility
by: Murphy, Brendan, et al.
Published: (2025) -
Exploiting Novel GPT-4 APIs
by: Pelrine, Kellin, et al.
Published: (2023) -
Complexity Matters: Effective Dimensionality as a Measure for Adversarial Robustness
by: Khachaturov, David, et al.
Published: (2024) -
Have You Poisoned My Data? Defending Neural Networks against Data Poisoning
by: De Gaspari, Fabio, et al.
Published: (2024)