LLMs Can Unlearn Refusal with Only 1,000 Benign Samples
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Guo, Yangyang, Xu, Ziwei, Liu, Si, Zheng, Zhiming, Kankanhalli, Mohan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Involuntary Jailbreak: On Self-Prompting Attacks
von: Guo, Yangyang, et al.
Veröffentlicht: (2025)
von: Guo, Yangyang, et al.
Veröffentlicht: (2025)
Technical Report for ICML 2024 TiFA Workshop MLLM Attack Challenge: Suffix Injection and Projected Gradient Descent Can Easily Fool An MLLM
von: Guo, Yangyang, et al.
Veröffentlicht: (2024)
von: Guo, Yangyang, et al.
Veröffentlicht: (2024)
DPTraj-PM: Differentially Private Trajectory Synthesis Using Prefix Tree and Markov Process
von: Wang, Nana, et al.
Veröffentlicht: (2024)
von: Wang, Nana, et al.
Veröffentlicht: (2024)
Do Prompts Guarantee Safety? Mitigating Toxicity from LLM Generations through Subspace Intervention
von: Singh, Himanshu, et al.
Veröffentlicht: (2026)
von: Singh, Himanshu, et al.
Veröffentlicht: (2026)
The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense
von: Guo, Yangyang, et al.
Veröffentlicht: (2024)
von: Guo, Yangyang, et al.
Veröffentlicht: (2024)
Privacy Risks and Preservation Methods in Explainable Artificial Intelligence: A Scoping Review
von: Allana, Sonal, et al.
Veröffentlicht: (2025)
von: Allana, Sonal, et al.
Veröffentlicht: (2025)
Releasing Malevolence from Benevolence: The Menace of Benign Data on Machine Unlearning
von: Ma, Binhao, et al.
Veröffentlicht: (2024)
von: Ma, Binhao, et al.
Veröffentlicht: (2024)
Really Unlearned? Verifying Machine Unlearning via Influential Sample Pairs
von: Xu, Heng, et al.
Veröffentlicht: (2024)
von: Xu, Heng, et al.
Veröffentlicht: (2024)
Forgetting Similar Samples: Can Machine Unlearning Do it Better?
von: Xu, Heng, et al.
Veröffentlicht: (2026)
von: Xu, Heng, et al.
Veröffentlicht: (2026)
Unified Neural Backdoor Removal with Only Few Clean Samples through Unlearning and Relearning
von: Min, Nay Myat, et al.
Veröffentlicht: (2024)
von: Min, Nay Myat, et al.
Veröffentlicht: (2024)
GESR: Graph-Based Edge Semantic Reconstruction for Stealthy Communication Detection with Benign-Only Training
von: Xu, Henghui, et al.
Veröffentlicht: (2026)
von: Xu, Henghui, et al.
Veröffentlicht: (2026)
Beyond Surface Alignment: Rebuilding LLMs Safety Mechanism via Probabilistically Ablating Refusal Direction
von: Xie, Yuanbo, et al.
Veröffentlicht: (2025)
von: Xie, Yuanbo, et al.
Veröffentlicht: (2025)
MCP Safety Training: Learning to Refuse Falsely Benign MCP Exploits using Improved Preference Alignment
von: Halloran, John
Veröffentlicht: (2025)
von: Halloran, John
Veröffentlicht: (2025)
Safety is Not Only About Refusal: Reasoning-Enhanced Fine-tuning for Interpretable LLM Safety
von: Zhang, Yuyou, et al.
Veröffentlicht: (2025)
von: Zhang, Yuyou, et al.
Veröffentlicht: (2025)
Refusal-Trained LLMs Are Easily Jailbroken As Browser Agents
von: Kumar, Priyanshu, et al.
Veröffentlicht: (2024)
von: Kumar, Priyanshu, et al.
Veröffentlicht: (2024)
Attack via Overfitting: 10-shot Benign Fine-tuning to Jailbreak LLMs
von: Xie, Zhixin, et al.
Veröffentlicht: (2025)
von: Xie, Zhixin, et al.
Veröffentlicht: (2025)
Self and Cross-Model Distillation for LLMs: Effective Methods for Refusal Pattern Alignment
von: Li, Jie, et al.
Veröffentlicht: (2024)
von: Li, Jie, et al.
Veröffentlicht: (2024)
Misrouter: Exploiting Routing Mechanisms for Input-Only Attacks on Mixture-of-Experts LLMs
von: Fei, Zekun, et al.
Veröffentlicht: (2026)
von: Fei, Zekun, et al.
Veröffentlicht: (2026)
Benign Fine-Tuning Breaks Safety Alignment in Audio LLMs
von: Roh, Jaechul, et al.
Veröffentlicht: (2026)
von: Roh, Jaechul, et al.
Veröffentlicht: (2026)
Split Unlearning
von: Yu, Guangsheng, et al.
Veröffentlicht: (2023)
von: Yu, Guangsheng, et al.
Veröffentlicht: (2023)
Do Reasoning LLMs Refuse What They Infer in Long Contexts?
von: Fu, Yu, et al.
Veröffentlicht: (2026)
von: Fu, Yu, et al.
Veröffentlicht: (2026)
Step-by-Step Reasoning Attack: Revealing 'Erased' Knowledge in Large Language Models
von: Sinha, Yash, et al.
Veröffentlicht: (2025)
von: Sinha, Yash, et al.
Veröffentlicht: (2025)
Self-Jailbreaking: Language Models Can Reason Themselves Out of Safety Alignment After Benign Reasoning Training
von: Yong, Zheng-Xin, et al.
Veröffentlicht: (2025)
von: Yong, Zheng-Xin, et al.
Veröffentlicht: (2025)
Unlearning Isn't Deletion: Investigating Reversibility of Machine Unlearning in LLMs
von: Xu, Xiaoyu, et al.
Veröffentlicht: (2025)
von: Xu, Xiaoyu, et al.
Veröffentlicht: (2025)
TRACE: Timely Retrieval and Alignment for Cybersecurity Knowledge Graph Construction and Expansion
von: Xu, Zijing, et al.
Veröffentlicht: (2026)
von: Xu, Zijing, et al.
Veröffentlicht: (2026)
Label Leakage Attacks in Machine Unlearning: A Parameter and Inversion-Based Approach
von: Zheng, Weidong, et al.
Veröffentlicht: (2026)
von: Zheng, Weidong, et al.
Veröffentlicht: (2026)
Few-Shot Truly Benign DPO Attack for Jailbreaking LLMs
von: Yoon, Sangyeon, et al.
Veröffentlicht: (2026)
von: Yoon, Sangyeon, et al.
Veröffentlicht: (2026)
TrojanPraise: Jailbreak LLMs via Benign Fine-Tuning
von: Xie, Zhixin, et al.
Veröffentlicht: (2026)
von: Xie, Zhixin, et al.
Veröffentlicht: (2026)
Privacy-Preserving Federated Unlearning with Certified Client Removal
von: Liu, Ziyao, et al.
Veröffentlicht: (2024)
von: Liu, Ziyao, et al.
Veröffentlicht: (2024)
Refusal Before Decoding: Detecting and Exploiting Refusal Signals in Intermediate LLM Activations
von: Collu, Matteo Gioele, et al.
Veröffentlicht: (2026)
von: Collu, Matteo Gioele, et al.
Veröffentlicht: (2026)
Unlearn and Burn: Adversarial Machine Unlearning Requests Destroy Model Accuracy
von: Huang, Yangsibo, et al.
Veröffentlicht: (2024)
von: Huang, Yangsibo, et al.
Veröffentlicht: (2024)
Learn What You Want to Unlearn: Unlearning Inversion Attacks against Machine Unlearning
von: Hu, Hongsheng, et al.
Veröffentlicht: (2024)
von: Hu, Hongsheng, et al.
Veröffentlicht: (2024)
Guaranteeing Data Privacy in Federated Unlearning with Dynamic User Participation
von: Liu, Ziyao, et al.
Veröffentlicht: (2024)
von: Liu, Ziyao, et al.
Veröffentlicht: (2024)
Towards Lifecycle Unlearning Commitment Management: Measuring Sample-level Approximate Unlearning Completeness
von: Wang, Cheng-Long, et al.
Veröffentlicht: (2024)
von: Wang, Cheng-Long, et al.
Veröffentlicht: (2024)
SAEs $\textit{Can}$ Improve Unlearning: Dynamic Sparse Autoencoder Guardrails for Precision Unlearning in LLMs
von: Muhamed, Aashiq, et al.
Veröffentlicht: (2025)
von: Muhamed, Aashiq, et al.
Veröffentlicht: (2025)
Why Do Aligned LLMs Remain Jailbreakable: Refusal-Escape Directions, Operator-Level Sources, and Safety-Utility Trade-off
von: Chen, Yu, et al.
Veröffentlicht: (2026)
von: Chen, Yu, et al.
Veröffentlicht: (2026)
CodeBC: A More Secure Large Language Model for Smart Contract Code Generation in Blockchain
von: Wang, Lingxiang, et al.
Veröffentlicht: (2025)
von: Wang, Lingxiang, et al.
Veröffentlicht: (2025)
Evaluating Disassembly Errors With Only Binaries
von: Wijayadi, Lambang Akbar, et al.
Veröffentlicht: (2025)
von: Wijayadi, Lambang Akbar, et al.
Veröffentlicht: (2025)
Machine Unlearning in Large Language Models
von: Chen, Kongyang, et al.
Veröffentlicht: (2024)
von: Chen, Kongyang, et al.
Veröffentlicht: (2024)
Can LLMs Classify CVEs? Investigating LLMs Capabilities in Computing CVSS Vectors
von: Marchiori, Francesco, et al.
Veröffentlicht: (2025)
von: Marchiori, Francesco, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Involuntary Jailbreak: On Self-Prompting Attacks
von: Guo, Yangyang, et al.
Veröffentlicht: (2025) -
Technical Report for ICML 2024 TiFA Workshop MLLM Attack Challenge: Suffix Injection and Projected Gradient Descent Can Easily Fool An MLLM
von: Guo, Yangyang, et al.
Veröffentlicht: (2024) -
DPTraj-PM: Differentially Private Trajectory Synthesis Using Prefix Tree and Markov Process
von: Wang, Nana, et al.
Veröffentlicht: (2024) -
Do Prompts Guarantee Safety? Mitigating Toxicity from LLM Generations through Subspace Intervention
von: Singh, Himanshu, et al.
Veröffentlicht: (2026) -
The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense
von: Guo, Yangyang, et al.
Veröffentlicht: (2024)