Poisoning Web-Scale Training Datasets is Practical
Fuente:
arXiv
Salvato in:
| Autori principali: | Carlini, Nicholas, Jagielski, Matthew, Choquette-Choo, Christopher A., Paleka, Daniel, Pearce, Will, Anderson, Hyrum, Terzis, Andreas, Thomas, Kurt, Tramèr, Florian |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Privacy Side Channels in Machine Learning Systems
di: Debenedetti, Edoardo, et al.
Pubblicazione: (2023)
di: Debenedetti, Edoardo, et al.
Pubblicazione: (2023)
LLMs unlock new paths to monetizing exploits
di: Carlini, Nicholas, et al.
Pubblicazione: (2025)
di: Carlini, Nicholas, et al.
Pubblicazione: (2025)
Auditing Private Prediction
di: Chadha, Karan, et al.
Pubblicazione: (2024)
di: Chadha, Karan, et al.
Pubblicazione: (2024)
Position: Considerations for Differentially Private Learning with Large-Scale Public Pretraining
di: Tramèr, Florian, et al.
Pubblicazione: (2022)
di: Tramèr, Florian, et al.
Pubblicazione: (2022)
Large-scale online deanonymization with LLMs
di: Lermen, Simon, et al.
Pubblicazione: (2026)
di: Lermen, Simon, et al.
Pubblicazione: (2026)
Persistent Pre-Training Poisoning of LLMs
di: Zhang, Yiming, et al.
Pubblicazione: (2024)
di: Zhang, Yiming, et al.
Pubblicazione: (2024)
The Last Iterate Advantage: Empirical Auditing and Principled Heuristic Analysis of Differentially Private SGD
di: Steinke, Thomas, et al.
Pubblicazione: (2024)
di: Steinke, Thomas, et al.
Pubblicazione: (2024)
Evading Black-box Classifiers Without Breaking Eggs
di: Debenedetti, Edoardo, et al.
Pubblicazione: (2023)
di: Debenedetti, Edoardo, et al.
Pubblicazione: (2023)
Privacy Ripple Effects from Adding or Removing Personal Information in Language Model Training
di: Borkar, Jaydeep, et al.
Pubblicazione: (2025)
di: Borkar, Jaydeep, et al.
Pubblicazione: (2025)
Adversarial Perturbations Cannot Reliably Protect Artists From Generative AI
di: Hönig, Robert, et al.
Pubblicazione: (2024)
di: Hönig, Robert, et al.
Pubblicazione: (2024)
Stealing Part of a Production Language Model
di: Carlini, Nicholas, et al.
Pubblicazione: (2024)
di: Carlini, Nicholas, et al.
Pubblicazione: (2024)
Are aligned neural networks adversarially aligned?
di: Carlini, Nicholas, et al.
Pubblicazione: (2023)
di: Carlini, Nicholas, et al.
Pubblicazione: (2023)
Exploring and Mitigating Adversarial Manipulation of Voting-Based Leaderboards
di: Huang, Yangsibo, et al.
Pubblicazione: (2025)
di: Huang, Yangsibo, et al.
Pubblicazione: (2025)
Adversarial ML Problems Are Getting Harder to Solve and to Evaluate
di: Rando, Javier, et al.
Pubblicazione: (2025)
di: Rando, Javier, et al.
Pubblicazione: (2025)
Defeating Prompt Injections by Design
di: Debenedetti, Edoardo, et al.
Pubblicazione: (2025)
di: Debenedetti, Edoardo, et al.
Pubblicazione: (2025)
Language Models May Verbatim Complete Text They Were Not Explicitly Trained On
di: Liu, Ken Ziyu, et al.
Pubblicazione: (2025)
di: Liu, Ken Ziyu, et al.
Pubblicazione: (2025)
Universal Jailbreak Backdoors from Poisoned Human Feedback
di: Rando, Javier, et al.
Pubblicazione: (2023)
di: Rando, Javier, et al.
Pubblicazione: (2023)
AutoAdvExBench: Benchmarking autonomous exploitation of adversarial example defenses
di: Carlini, Nicholas, et al.
Pubblicazione: (2025)
di: Carlini, Nicholas, et al.
Pubblicazione: (2025)
Near Exact Privacy Amplification for Matrix Mechanisms
di: Choquette-Choo, Christopher A., et al.
Pubblicazione: (2024)
di: Choquette-Choo, Christopher A., et al.
Pubblicazione: (2024)
Optimal Rates for $O(1)$-Smooth DP-SCO with a Single Epoch and Large Batches
di: Choquette-Choo, Christopher A., et al.
Pubblicazione: (2024)
di: Choquette-Choo, Christopher A., et al.
Pubblicazione: (2024)
Privacy Amplification for Matrix Mechanisms
di: Choquette-Choo, Christopher A., et al.
Pubblicazione: (2023)
di: Choquette-Choo, Christopher A., et al.
Pubblicazione: (2023)
Evaluating the Robustness of a Production Malware Detection System to Transferable Adversarial Attacks
di: Nasr, Milad, et al.
Pubblicazione: (2025)
di: Nasr, Milad, et al.
Pubblicazione: (2025)
Evaluating the Robustness of the "Ensemble Everything Everywhere" Defense
di: Zhang, Jie, et al.
Pubblicazione: (2024)
di: Zhang, Jie, et al.
Pubblicazione: (2024)
Query-Based Adversarial Prompt Generation
di: Hayase, Jonathan, et al.
Pubblicazione: (2024)
di: Hayase, Jonathan, et al.
Pubblicazione: (2024)
The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against Llm Jailbreaks and Prompt Injections
di: Nasr, Milad, et al.
Pubblicazione: (2025)
di: Nasr, Milad, et al.
Pubblicazione: (2025)
Phantom: General Backdoor Attacks on Retrieval Augmented Language Generation
di: Chaudhari, Harsh, et al.
Pubblicazione: (2024)
di: Chaudhari, Harsh, et al.
Pubblicazione: (2024)
Privacy Backdoors: Stealing Data with Corrupted Pretrained Models
di: Feng, Shanglun, et al.
Pubblicazione: (2024)
di: Feng, Shanglun, et al.
Pubblicazione: (2024)
Covert Attacks on Machine Learning Training in Passively Secure MPC
di: Jagielski, Matthew, et al.
Pubblicazione: (2025)
di: Jagielski, Matthew, et al.
Pubblicazione: (2025)
Membership Inference Attacks Cannot Prove that a Model Was Trained On Your Data
di: Zhang, Jie, et al.
Pubblicazione: (2024)
di: Zhang, Jie, et al.
Pubblicazione: (2024)
Thought-Transfer: Indirect Targeted Poisoning Attacks on Chain-of-Thought Reasoning Models
di: Chaudhari, Harsh, et al.
Pubblicazione: (2026)
di: Chaudhari, Harsh, et al.
Pubblicazione: (2026)
Certified Robustness to Clean-Label Poisoning Using Diffusion Denoising
di: Hong, Sanghyun, et al.
Pubblicazione: (2024)
di: Hong, Sanghyun, et al.
Pubblicazione: (2024)
Cutting through buggy adversarial example defenses: fixing 1 line of code breaks Sabre
di: Carlini, Nicholas
Pubblicazione: (2024)
di: Carlini, Nicholas
Pubblicazione: (2024)
Privacy Backdoors: Enhancing Membership Inference through Poisoning Pre-trained Models
di: Wen, Yuxin, et al.
Pubblicazione: (2024)
di: Wen, Yuxin, et al.
Pubblicazione: (2024)
Adversarial Search Engine Optimization for Large Language Models
di: Nestaas, Fredrik, et al.
Pubblicazione: (2024)
di: Nestaas, Fredrik, et al.
Pubblicazione: (2024)
Evaluations of Machine Learning Privacy Defenses are Misleading
di: Aerni, Michael, et al.
Pubblicazione: (2024)
di: Aerni, Michael, et al.
Pubblicazione: (2024)
Lessons from Defending Gemini Against Indirect Prompt Injections
di: Shi, Chongyang, et al.
Pubblicazione: (2025)
di: Shi, Chongyang, et al.
Pubblicazione: (2025)
Mellivora Capensis: A Backdoor-Free Training Framework on the Poisoned Dataset without Auxiliary Data
di: Pu, Yuwen, et al.
Pubblicazione: (2024)
di: Pu, Yuwen, et al.
Pubblicazione: (2024)
On Evaluating the Durability of Safeguards for Open-Weight LLMs
di: Qi, Xiangyu, et al.
Pubblicazione: (2024)
di: Qi, Xiangyu, et al.
Pubblicazione: (2024)
EMBER2024 -- A Benchmark Dataset for Holistic Evaluation of Malware Classifiers
di: Joyce, Robert J., et al.
Pubblicazione: (2025)
di: Joyce, Robert J., et al.
Pubblicazione: (2025)
Privacy Auditing of Large Language Models
di: Panda, Ashwinee, et al.
Pubblicazione: (2025)
di: Panda, Ashwinee, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Privacy Side Channels in Machine Learning Systems
di: Debenedetti, Edoardo, et al.
Pubblicazione: (2023) -
LLMs unlock new paths to monetizing exploits
di: Carlini, Nicholas, et al.
Pubblicazione: (2025) -
Auditing Private Prediction
di: Chadha, Karan, et al.
Pubblicazione: (2024) -
Position: Considerations for Differentially Private Learning with Large-Scale Public Pretraining
di: Tramèr, Florian, et al.
Pubblicazione: (2022) -
Large-scale online deanonymization with LLMs
di: Lermen, Simon, et al.
Pubblicazione: (2026)