Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods
Fuente:
arXiv
Saved in:
| Main Authors: | Jang, Yeonwoo, Hossain, Shariqah, Sreevatsa, Ashwin, Cruz, Diogo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Machine Unlearning Fails to Remove Data Poisoning Attacks
by: Pawelczyk, Martin, et al.
Published: (2024)
by: Pawelczyk, Martin, et al.
Published: (2024)
Superficial Safety Alignment Hypothesis
by: Li, Jianwei, et al.
Published: (2024)
by: Li, Jianwei, et al.
Published: (2024)
Hidden Poison: Machine Unlearning Enables Camouflaged Poisoning Attacks
by: Di, Jimmy Z., et al.
Published: (2022)
by: Di, Jimmy Z., et al.
Published: (2022)
Harry Potter is Still Here! Probing Knowledge Leakage in Targeted Unlearned Large Language Models via Automated Adversarial Prompting
by: To, Bang Trinh Tran, et al.
Published: (2025)
by: To, Bang Trinh Tran, et al.
Published: (2025)
Towards Robust Knowledge Unlearning: An Adversarial Framework for Assessing and Improving Unlearning Robustness in Large Language Models
by: Yuan, Hongbang, et al.
Published: (2024)
by: Yuan, Hongbang, et al.
Published: (2024)
Formalizing and Benchmarking Prompt Injection Attacks and Defenses
by: Liu, Yupei, et al.
Published: (2023)
by: Liu, Yupei, et al.
Published: (2023)
Safety Alignment Can Be Not Superficial With Explicit Safety Signals
by: Li, Jianwei, et al.
Published: (2025)
by: Li, Jianwei, et al.
Published: (2025)
Textual Unlearning Gives a False Sense of Unlearning
by: Du, Jiacheng, et al.
Published: (2024)
by: Du, Jiacheng, et al.
Published: (2024)
SequentialBreak: Large Language Models Can be Fooled by Embedding Jailbreak Prompts into Sequential Prompt Chains
by: Saiem, Bijoy Ahmed, et al.
Published: (2024)
by: Saiem, Bijoy Ahmed, et al.
Published: (2024)
Revealing Weaknesses in Text Watermarking Through Self-Information Rewrite Attacks
by: Cheng, Yixin, et al.
Published: (2025)
by: Cheng, Yixin, et al.
Published: (2025)
Enhancing Prompt Injection Attacks to LLMs via Poisoning Alignment
by: Shao, Zedian, et al.
Published: (2024)
by: Shao, Zedian, et al.
Published: (2024)
Unlearned but Not Forgotten: Data Extraction after Exact Unlearning in LLM
by: Wu, Xiaoyu, et al.
Published: (2025)
by: Wu, Xiaoyu, et al.
Published: (2025)
Unlearning Isn't Deletion: Investigating Reversibility of Machine Unlearning in LLMs
by: Xu, Xiaoyu, et al.
Published: (2025)
by: Xu, Xiaoyu, et al.
Published: (2025)
Knowledge Distillation-Based Model Extraction Attack using GAN-based Private Counterfactual Explanations
by: Ezzeddine, Fatima, et al.
Published: (2024)
by: Ezzeddine, Fatima, et al.
Published: (2024)
UnUnlearning: Unlearning is not sufficient for content regulation in advanced generative AI
by: Shumailov, Ilia, et al.
Published: (2024)
by: Shumailov, Ilia, et al.
Published: (2024)
Measuring Real-World Prompt Injection Attacks in LLM-based Resume Screening
by: Zhang, Mohan, et al.
Published: (2026)
by: Zhang, Mohan, et al.
Published: (2026)
Checkpoint-GCG: Auditing and Attacking Fine-Tuning-Based Prompt Injection Defenses
by: Yang, Xiaoxue, et al.
Published: (2025)
by: Yang, Xiaoxue, et al.
Published: (2025)
SAEs $\textit{Can}$ Improve Unlearning: Dynamic Sparse Autoencoder Guardrails for Precision Unlearning in LLMs
by: Muhamed, Aashiq, et al.
Published: (2025)
by: Muhamed, Aashiq, et al.
Published: (2025)
KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs
by: Liang, Buyun, et al.
Published: (2025)
by: Liang, Buyun, et al.
Published: (2025)
Misaligned Roles, Misplaced Images: Structural Input Perturbations Expose Multimodal Alignment Blind Spots
by: Shayegani, Erfan, et al.
Published: (2025)
by: Shayegani, Erfan, et al.
Published: (2025)
Just Do It!? Computer-Use Agents Exhibit Blind Goal-Directedness
by: Shayegani, Erfan, et al.
Published: (2025)
by: Shayegani, Erfan, et al.
Published: (2025)
Urania: Differentially Private Insights into AI Use
by: Liu, Daogao, et al.
Published: (2025)
by: Liu, Daogao, et al.
Published: (2025)
Clio: Privacy-Preserving Insights into Real-World AI Use
by: Tamkin, Alex, et al.
Published: (2024)
by: Tamkin, Alex, et al.
Published: (2024)
Can a large language model be a gaslighter?
by: Li, Wei, et al.
Published: (2024)
by: Li, Wei, et al.
Published: (2024)
An In-Depth Investigation of Data Collection in LLM App Ecosystems
by: Wu, Yuhao, et al.
Published: (2024)
by: Wu, Yuhao, et al.
Published: (2024)
What Makes an Evaluation Useful? Common Pitfalls and Best Practices
by: Gekker, Gil, et al.
Published: (2025)
by: Gekker, Gil, et al.
Published: (2025)
IsolateGPT: An Execution Isolation Architecture for LLM-Based Agentic Systems
by: Wu, Yuhao, et al.
Published: (2024)
by: Wu, Yuhao, et al.
Published: (2024)
Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
by: Zhang, Andy K., et al.
Published: (2024)
by: Zhang, Andy K., et al.
Published: (2024)
LLM Platform Security: Applying a Systematic Evaluation Framework to OpenAI's ChatGPT Plugins
by: Iqbal, Umar, et al.
Published: (2023)
by: Iqbal, Umar, et al.
Published: (2023)
Generative AI Security: Challenges and Countermeasures
by: Zhu, Banghua, et al.
Published: (2024)
by: Zhu, Banghua, et al.
Published: (2024)
An Adversarial Perspective on Machine Unlearning for AI Safety
by: Łucki, Jakub, et al.
Published: (2024)
by: Łucki, Jakub, et al.
Published: (2024)
Trustless Audits without Revealing Data or Models
by: Waiwitlikhit, Suppakit, et al.
Published: (2024)
by: Waiwitlikhit, Suppakit, et al.
Published: (2024)
Unlearning Backdoor Attacks for LLMs with Weak-to-Strong Knowledge Distillation
by: Zhao, Shuai, et al.
Published: (2024)
by: Zhao, Shuai, et al.
Published: (2024)
Machine Unlearning of Pre-trained Large Language Models
by: Yao, Jin, et al.
Published: (2024)
by: Yao, Jin, et al.
Published: (2024)
RECAP: A Resource-Efficient Method for Adversarial Prompting in Large Language Models
by: Chugh, Rishit
Published: (2026)
by: Chugh, Rishit
Published: (2026)
LLM Ghostbusters: Surgical Hallucination Suppression via Adaptive Unlearning
by: Spracklen, Joseph, et al.
Published: (2026)
by: Spracklen, Joseph, et al.
Published: (2026)
FIT to Forget: Robust Continual Unlearning for Large Language Models
by: Xu, Xiaoyu, et al.
Published: (2026)
by: Xu, Xiaoyu, et al.
Published: (2026)
OBLIVIATE: Robust and Practical Machine Unlearning for Large Language Models
by: Xu, Xiaoyu, et al.
Published: (2025)
by: Xu, Xiaoyu, et al.
Published: (2025)
From Domains to Instances: Dual-Granularity Data Synthesis for LLM Unlearning
by: Xu, Xiaoyu, et al.
Published: (2026)
by: Xu, Xiaoyu, et al.
Published: (2026)
A Survey of Privacy-Preserving Model Explanations: Privacy Risks, Attacks, and Countermeasures
by: Nguyen, Thanh Tam, et al.
Published: (2024)
by: Nguyen, Thanh Tam, et al.
Published: (2024)
Similar Items
-
Machine Unlearning Fails to Remove Data Poisoning Attacks
by: Pawelczyk, Martin, et al.
Published: (2024) -
Superficial Safety Alignment Hypothesis
by: Li, Jianwei, et al.
Published: (2024) -
Hidden Poison: Machine Unlearning Enables Camouflaged Poisoning Attacks
by: Di, Jimmy Z., et al.
Published: (2022) -
Harry Potter is Still Here! Probing Knowledge Leakage in Targeted Unlearned Large Language Models via Automated Adversarial Prompting
by: To, Bang Trinh Tran, et al.
Published: (2025) -
Towards Robust Knowledge Unlearning: An Adversarial Framework for Assessing and Improving Unlearning Robustness in Large Language Models
by: Yuan, Hongbang, et al.
Published: (2024)