Poisoning Attacks on LLMs Require a Near-constant Number of Poison Samples
Fuente:
arXiv
Salvato in:
| Autori principali: | Souly, Alexandra, Rando, Javier, Chapman, Ed, Davies, Xander, Hasircioglu, Burak, Shereen, Ezzeldin, Mougan, Carlos, Mavroudis, Vasilios, Jones, Erik, Hicks, Chris, Carlini, Nicholas, Gal, Yarin, Kirk, Robert |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
One Pic is All it Takes: Poisoning Visual Document Retrieval Augmented Generation with a Single Image
di: Shereen, Ezzeldin, et al.
Pubblicazione: (2025)
di: Shereen, Ezzeldin, et al.
Pubblicazione: (2025)
Beyond Training-time Poisoning: Component-level and Post-training Backdoors in Deep Reinforcement Learning
di: Vyas, Sanyam, et al.
Pubblicazione: (2025)
di: Vyas, Sanyam, et al.
Pubblicazione: (2025)
From Promise to Peril: Rethinking Cybersecurity Red and Blue Teaming in the Age of LLMs
di: Abuadbba, Alsharif, et al.
Pubblicazione: (2025)
di: Abuadbba, Alsharif, et al.
Pubblicazione: (2025)
Direction for Detection: A Survey of Automated Vulnerability Detection and all of its Pain Points
di: Ristea, Dan, et al.
Pubblicazione: (2024)
di: Ristea, Dan, et al.
Pubblicazione: (2024)
Fundamental Limitations in Pointwise Defences of LLM Finetuning APIs
di: Davies, Xander, et al.
Pubblicazione: (2025)
di: Davies, Xander, et al.
Pubblicazione: (2025)
Reasoning Introduces New Poisoning Attacks Yet Makes Them More Complicated
di: Foerster, Hanna, et al.
Pubblicazione: (2025)
di: Foerster, Hanna, et al.
Pubblicazione: (2025)
On Efficient Bayesian Exploration in Model-Based Reinforcement Learning
di: Caron, Alberto, et al.
Pubblicazione: (2025)
di: Caron, Alberto, et al.
Pubblicazione: (2025)
Towards Causal Model-Based Policy Optimization
di: Caron, Alberto, et al.
Pubblicazione: (2025)
di: Caron, Alberto, et al.
Pubblicazione: (2025)
A View on Out-of-Distribution Identification from a Statistical Testing Theory Perspective
di: Caron, Alberto, et al.
Pubblicazione: (2024)
di: Caron, Alberto, et al.
Pubblicazione: (2024)
Nearest Neighbour with Bandit Feedback
di: Pasteris, Stephen, et al.
Pubblicazione: (2023)
di: Pasteris, Stephen, et al.
Pubblicazione: (2023)
Mitigating Deep Reinforcement Learning Backdoors in the Neural Activation Space
di: Vyas, Sanyam, et al.
Pubblicazione: (2024)
di: Vyas, Sanyam, et al.
Pubblicazione: (2024)
Beyond Rewards in Reinforcement Learning for Cyber Defence
di: Bates, Elizabeth, et al.
Pubblicazione: (2026)
di: Bates, Elizabeth, et al.
Pubblicazione: (2026)
Fairness with Exponential Weights
di: Pasteris, Stephen, et al.
Pubblicazione: (2024)
di: Pasteris, Stephen, et al.
Pubblicazione: (2024)
Less is more? Rewards in RL for Cyber Defence
di: Bates, Elizabeth, et al.
Pubblicazione: (2025)
di: Bates, Elizabeth, et al.
Pubblicazione: (2025)
Extraction Propagation
di: Pasteris, Stephen, et al.
Pubblicazione: (2024)
di: Pasteris, Stephen, et al.
Pubblicazione: (2024)
Distributed Detection of Adversarial Attacks in Multi-Agent Reinforcement Learning with Continuous Action Space
di: Kazari, Kiarash, et al.
Pubblicazione: (2025)
di: Kazari, Kiarash, et al.
Pubblicazione: (2025)
Zero-Trust Network Access (ZTNA)
di: Mavroudis, Vasilios
Pubblicazione: (2024)
di: Mavroudis, Vasilios
Pubblicazione: (2024)
Persistent Pre-Training Poisoning of LLMs
di: Zhang, Yiming, et al.
Pubblicazione: (2024)
di: Zhang, Yiming, et al.
Pubblicazione: (2024)
Evaluating whether AI models would sabotage AI safety research
di: Kirk, Robert, et al.
Pubblicazione: (2026)
di: Kirk, Robert, et al.
Pubblicazione: (2026)
UK AISI Alignment Evaluation Case-Study
di: Souly, Alexandra, et al.
Pubblicazione: (2026)
di: Souly, Alexandra, et al.
Pubblicazione: (2026)
Universal Jailbreak Backdoors from Poisoned Human Feedback
di: Rando, Javier, et al.
Pubblicazione: (2023)
di: Rando, Javier, et al.
Pubblicazione: (2023)
Environment Complexity and Nash Equilibria in a Sequential Social Dilemma
di: Yasir, Mustafa, et al.
Pubblicazione: (2024)
di: Yasir, Mustafa, et al.
Pubblicazione: (2024)
Autonomous Network Defence using Reinforcement Learning
di: Foley, Myles, et al.
Pubblicazione: (2024)
di: Foley, Myles, et al.
Pubblicazione: (2024)
Online Convex Optimisation: The Optimal Switching Regret for all Segmentations Simultaneously
di: Pasteris, Stephen, et al.
Pubblicazione: (2024)
di: Pasteris, Stephen, et al.
Pubblicazione: (2024)
Inherently Interpretable and Uncertainty-Aware Models for Online Learning in Cyber-Security Problems
di: Kolicic, Benjamin, et al.
Pubblicazione: (2024)
di: Kolicic, Benjamin, et al.
Pubblicazione: (2024)
CybORG++: An Enhanced Gym for the Development of Autonomous Cyber Agents
di: Emerson, Harry, et al.
Pubblicazione: (2024)
di: Emerson, Harry, et al.
Pubblicazione: (2024)
Entity-based Reinforcement Learning for Autonomous Cyber Defence
di: Thompson, Isaac Symes, et al.
Pubblicazione: (2024)
di: Thompson, Isaac Symes, et al.
Pubblicazione: (2024)
HonestCyberEval: An AI Cyber Risk Benchmark for Automated Software Exploitation
di: Ristea, Dan, et al.
Pubblicazione: (2024)
di: Ristea, Dan, et al.
Pubblicazione: (2024)
Analysis of Publicly Accessible Operational Technology and Associated Risks
di: Rodda, Matthew, et al.
Pubblicazione: (2025)
di: Rodda, Matthew, et al.
Pubblicazione: (2025)
Quantifying Mix Network Privacy Erosion with Generative Models
di: Mavroudis, Vasilios, et al.
Pubblicazione: (2025)
di: Mavroudis, Vasilios, et al.
Pubblicazione: (2025)
Referential Security as a New Paradigm for AI Evaluations
di: Ristea, Dan, et al.
Pubblicazione: (2026)
di: Ristea, Dan, et al.
Pubblicazione: (2026)
Hidden Poison: Machine Unlearning Enables Camouflaged Poisoning Attacks
di: Di, Jimmy Z., et al.
Pubblicazione: (2022)
di: Di, Jimmy Z., et al.
Pubblicazione: (2022)
PoisonCatcher: Revealing and Identifying LDP Poisoning Attacks in IIoT
di: Shuai, Lisha, et al.
Pubblicazione: (2024)
di: Shuai, Lisha, et al.
Pubblicazione: (2024)
What if we could hot swap our Biometrics?
di: Crowcroft, Jon, et al.
Pubblicazione: (2025)
di: Crowcroft, Jon, et al.
Pubblicazione: (2025)
An Attentive Graph Agent for Topology-Adaptive Cyber Defence
di: Sandoval, Ilya Orson, et al.
Pubblicazione: (2025)
di: Sandoval, Ilya Orson, et al.
Pubblicazione: (2025)
Boundary Point Jailbreaking of Black-Box LLMs
di: Davies, Xander, et al.
Pubblicazione: (2026)
di: Davies, Xander, et al.
Pubblicazione: (2026)
Certified Robustness to Clean-Label Poisoning Using Diffusion Denoising
di: Hong, Sanghyun, et al.
Pubblicazione: (2024)
di: Hong, Sanghyun, et al.
Pubblicazione: (2024)
PoisonArena: Uncovering Competing Poisoning Attacks in Retrieval-Augmented Generation
di: Chen, Liuji, et al.
Pubblicazione: (2025)
di: Chen, Liuji, et al.
Pubblicazione: (2025)
Poisoned-MRAG: Knowledge Poisoning Attacks to Multimodal Retrieval Augmented Generation
di: Liu, Yinuo, et al.
Pubblicazione: (2025)
di: Liu, Yinuo, et al.
Pubblicazione: (2025)
Transferable Availability Poisoning Attacks
di: Liu, Yiyong, et al.
Pubblicazione: (2023)
di: Liu, Yiyong, et al.
Pubblicazione: (2023)
Documenti analoghi
-
One Pic is All it Takes: Poisoning Visual Document Retrieval Augmented Generation with a Single Image
di: Shereen, Ezzeldin, et al.
Pubblicazione: (2025) -
Beyond Training-time Poisoning: Component-level and Post-training Backdoors in Deep Reinforcement Learning
di: Vyas, Sanyam, et al.
Pubblicazione: (2025) -
From Promise to Peril: Rethinking Cybersecurity Red and Blue Teaming in the Age of LLMs
di: Abuadbba, Alsharif, et al.
Pubblicazione: (2025) -
Direction for Detection: A Survey of Automated Vulnerability Detection and all of its Pain Points
di: Ristea, Dan, et al.
Pubblicazione: (2024) -
Fundamental Limitations in Pointwise Defences of LLM Finetuning APIs
di: Davies, Xander, et al.
Pubblicazione: (2025)