UnUnlearning: Unlearning is not sufficient for content regulation in advanced generative AI
Fuente:
arXiv
Saved in:
| Main Authors: | Shumailov, Ilia, Hayes, Jamie, Triantafillou, Eleni, Ortiz-Jimenez, Guillermo, Papernot, Nicolas, Jagielski, Matthew, Yona, Itay, Howard, Heidi, Bagdasaryan, Eugene |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Inexact Unlearning Needs More Careful Evaluations to Avoid a False Sense of Privacy
by: Hayes, Jamie, et al.
Published: (2024)
by: Hayes, Jamie, et al.
Published: (2024)
Buffer Overflow in Mixture of Experts
by: Hayes, Jamie, et al.
Published: (2024)
by: Hayes, Jamie, et al.
Published: (2024)
Stealing User Prompts from Mixture of Experts
by: Yona, Itay, et al.
Published: (2024)
by: Yona, Itay, et al.
Published: (2024)
Interpreting the Repeated Token Phenomenon in Large Language Models
by: Yona, Itay, et al.
Published: (2025)
by: Yona, Itay, et al.
Published: (2025)
Fairness Feedback Loops: Training on Synthetic Data Amplifies Bias
by: Wyllie, Sierra, et al.
Published: (2024)
by: Wyllie, Sierra, et al.
Published: (2024)
Leveraging Per-Instance Privacy for Machine Unlearning
by: Sepahvand, Nazanin Mohammadi, et al.
Published: (2025)
by: Sepahvand, Nazanin Mohammadi, et al.
Published: (2025)
Extracting alignment data in open models
by: Barbero, Federico, et al.
Published: (2025)
by: Barbero, Federico, et al.
Published: (2025)
Data Selection for Transfer Unlearning
by: Sepahvand, Nazanin Mohammadi, et al.
Published: (2024)
by: Sepahvand, Nazanin Mohammadi, et al.
Published: (2024)
Cascading Adversarial Bias from Injection to Distillation in Language Models
by: Chaudhari, Harsh, et al.
Published: (2025)
by: Chaudhari, Harsh, et al.
Published: (2025)
Beyond Labeling Oracles: What does it mean to steal ML models?
by: Shafran, Avital, et al.
Published: (2023)
by: Shafran, Avital, et al.
Published: (2023)
Benchmarking Unlearning for Vision Transformers
by: Zhao, Kairan, et al.
Published: (2026)
by: Zhao, Kairan, et al.
Published: (2026)
Architectural Neural Backdoors from First Principles
by: Langford, Harry, et al.
Published: (2024)
by: Langford, Harry, et al.
Published: (2024)
Beyond Laplace and Gaussian: Exploring the Generalized Gaussian Mechanism for Private Machine Learning
by: Rinberg, Roy, et al.
Published: (2025)
by: Rinberg, Roy, et al.
Published: (2025)
Breach By A Thousand Leaks: Unsafe Information Leakage in `Safe' AI Responses
by: Glukhov, David, et al.
Published: (2024)
by: Glukhov, David, et al.
Published: (2024)
Gradients Look Alike: Sensitivity is Often Overestimated in DP-SGD
by: Thudi, Anvith, et al.
Published: (2023)
by: Thudi, Anvith, et al.
Published: (2023)
When Vision Fails: Text Attacks Against ViT and OCR
by: Boucher, Nicholas, et al.
Published: (2023)
by: Boucher, Nicholas, et al.
Published: (2023)
Improved Localized Machine Unlearning Through the Lens of Memorization
by: Torkzadehmahani, Reihaneh, et al.
Published: (2024)
by: Torkzadehmahani, Reihaneh, et al.
Published: (2024)
Beyond Slow Signs in High-fidelity Model Extraction
by: Foerster, Hanna, et al.
Published: (2024)
by: Foerster, Hanna, et al.
Published: (2024)
Measuring memorization in RLHF for code completion
by: Pappu, Aneesh, et al.
Published: (2024)
by: Pappu, Aneesh, et al.
Published: (2024)
Operationalizing Contextual Integrity in Privacy-Conscious Assistants
by: Ghalebikesabi, Sahra, et al.
Published: (2024)
by: Ghalebikesabi, Sahra, et al.
Published: (2024)
Verifiable and Provably Secure Machine Unlearning
by: Eisenhofer, Thorsten, et al.
Published: (2022)
by: Eisenhofer, Thorsten, et al.
Published: (2022)
Thought-Transfer: Indirect Targeted Poisoning Attacks on Chain-of-Thought Reasoning Models
by: Chaudhari, Harsh, et al.
Published: (2026)
by: Chaudhari, Harsh, et al.
Published: (2026)
Fast Exact Unlearning for In-Context Learning Data for LLMs
by: Muresanu, Andrei I., et al.
Published: (2024)
by: Muresanu, Andrei I., et al.
Published: (2024)
Measuring memorization in language models via probabilistic extraction
by: Hayes, Jamie, et al.
Published: (2024)
by: Hayes, Jamie, et al.
Published: (2024)
The Curse of Recursion: Training on Generated Data Makes Models Forget
by: Shumailov, Ilia, et al.
Published: (2023)
by: Shumailov, Ilia, et al.
Published: (2023)
Machine Unlearning Doesn't Do What You Think: Lessons for Generative AI Policy and Research
by: Cooper, A. Feder, et al.
Published: (2024)
by: Cooper, A. Feder, et al.
Published: (2024)
From Dormant to Deleted: Tamper-Resistant Unlearning Through Weight-Space Regularization
by: Siddiqui, Shoaib Ahmed, et al.
Published: (2025)
by: Siddiqui, Shoaib Ahmed, et al.
Published: (2025)
Gauss-Newton Unlearning for the LLM Era
by: McKinney, Lev, et al.
Published: (2026)
by: McKinney, Lev, et al.
Published: (2026)
To Each (Textual Sequence) Its Own: Improving Memorized-Data Unlearning in Large Language Models
by: Barbulescu, George-Octavian, et al.
Published: (2024)
by: Barbulescu, George-Octavian, et al.
Published: (2024)
Soft Instruction De-escalation Defense
by: Walter, Nils Philipp, et al.
Published: (2025)
by: Walter, Nils Philipp, et al.
Published: (2025)
You Don't Need All That Attention: Surgical Memorization Mitigation in Text-to-Image Diffusion Models
by: Zhao, Kairan, et al.
Published: (2026)
by: Zhao, Kairan, et al.
Published: (2026)
Large Language Models Can Verbatim Reproduce Long Malicious Sequences
by: Lin, Sharon, et al.
Published: (2025)
by: Lin, Sharon, et al.
Published: (2025)
Quantamination: Dynamic Quantization Leaks Your Data Across the Batch
by: Foerster, Hanna, et al.
Published: (2026)
by: Foerster, Hanna, et al.
Published: (2026)
Auditing Private Prediction
by: Chadha, Karan, et al.
Published: (2024)
by: Chadha, Karan, et al.
Published: (2024)
Reasoning Introduces New Poisoning Attacks Yet Makes Them More Complicated
by: Foerster, Hanna, et al.
Published: (2025)
by: Foerster, Hanna, et al.
Published: (2025)
Lessons from Defending Gemini Against Indirect Prompt Injections
by: Shi, Chongyang, et al.
Published: (2025)
by: Shi, Chongyang, et al.
Published: (2025)
The sufficiently trapped surface
by: Kontou, Eleni-Alexandra
Published: (2026)
by: Kontou, Eleni-Alexandra
Published: (2026)
Adversarial Illusions in Multi-Modal Embeddings
by: Zhang, Tingwei, et al.
Published: (2023)
by: Zhang, Tingwei, et al.
Published: (2023)
CaMeLs Can Use Computers Too: System-level Security for Computer Use Agents
by: Foerster, Hanna, et al.
Published: (2026)
by: Foerster, Hanna, et al.
Published: (2026)
Architectural Backdoors for Within-Batch Data Stealing and Model Inference Manipulation
by: Küchler, Nicolas, et al.
Published: (2025)
by: Küchler, Nicolas, et al.
Published: (2025)
Similar Items
-
Inexact Unlearning Needs More Careful Evaluations to Avoid a False Sense of Privacy
by: Hayes, Jamie, et al.
Published: (2024) -
Buffer Overflow in Mixture of Experts
by: Hayes, Jamie, et al.
Published: (2024) -
Stealing User Prompts from Mixture of Experts
by: Yona, Itay, et al.
Published: (2024) -
Interpreting the Repeated Token Phenomenon in Large Language Models
by: Yona, Itay, et al.
Published: (2025) -
Fairness Feedback Loops: Training on Synthetic Data Amplifies Bias
by: Wyllie, Sierra, et al.
Published: (2024)