Releasing Malevolence from Benevolence: The Menace of Benign Data on Machine Unlearning
Fuente:
arXiv
Saved in:
| Main Authors: | Ma, Binhao, Zheng, Tianhang, Hu, Hongsheng, Wang, Di, Wang, Shuo, Ba, Zhongjie, Qin, Zhan, Ren, Kui |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learn What You Want to Unlearn: Unlearning Inversion Attacks against Machine Unlearning
by: Hu, Hongsheng, et al.
Published: (2024)
by: Hu, Hongsheng, et al.
Published: (2024)
MAJIC: Markovian Adaptive Jailbreaking via Iterative Composition of Diverse Innovative Strategies
by: Qi, Weiwei, et al.
Published: (2025)
by: Qi, Weiwei, et al.
Published: (2025)
Untargeted Jailbreak Attack
by: Huang, Xinzhe, et al.
Published: (2025)
by: Huang, Xinzhe, et al.
Published: (2025)
ERASER: Machine Unlearning in MLaaS via an Inference Serving-Aware Approach
by: Hu, Yuke, et al.
Published: (2023)
by: Hu, Yuke, et al.
Published: (2023)
WMCopier: Forging Invisible Image Watermarks on Arbitrary Images
by: Dong, Ziping, et al.
Published: (2025)
by: Dong, Ziping, et al.
Published: (2025)
Towards Identification and Intervention of Safety-Critical Parameters in Large Language Models
by: Qi, Weiwei, et al.
Published: (2026)
by: Qi, Weiwei, et al.
Published: (2026)
A Duty to Forget, a Right to be Assured? Exposing Vulnerabilities in Machine Unlearning Services
by: Hu, Hongsheng, et al.
Published: (2023)
by: Hu, Hongsheng, et al.
Published: (2023)
Attack-Resistant Watermarking for AIGC Image Forensics via Diffusion-based Semantic Deflection
by: Liu, Qingyu, et al.
Published: (2026)
by: Liu, Qingyu, et al.
Published: (2026)
"Training robust watermarking model may hurt authentication!'' Exploring and Mitigating the Identity Leakage in Robust Watermarking
by: Zhang, Xinyu, et al.
Published: (2026)
by: Zhang, Xinyu, et al.
Published: (2026)
SurrogatePrompt: Bypassing the Safety Filter of Text-to-Image Models via Substitution
by: Ba, Zhongjie, et al.
Published: (2023)
by: Ba, Zhongjie, et al.
Published: (2023)
BadFU: Backdoor Federated Learning through Adversarial Machine Unlearning
by: Lu, Bingguang, et al.
Published: (2025)
by: Lu, Bingguang, et al.
Published: (2025)
TRACE: Task-Aware Adaptive Self-Evolving Agentic Jailbreaking
by: Zeng, Churui, et al.
Published: (2026)
by: Zeng, Churui, et al.
Published: (2026)
Dynamic Target Attack
by: Xiu, Kedong, et al.
Published: (2025)
by: Xiu, Kedong, et al.
Published: (2025)
SWAT: A System-Wide Approach to Tunable Leakage Mitigation in Encrypted Data Stores
by: Zheng, Leqian, et al.
Published: (2023)
by: Zheng, Leqian, et al.
Published: (2023)
Towards Label-Only Membership Inference Attack against Pre-trained Large Language Models
by: He, Yu, et al.
Published: (2025)
by: He, Yu, et al.
Published: (2025)
LLMs Can Unlearn Refusal with Only 1,000 Benign Samples
by: Guo, Yangyang, et al.
Published: (2026)
by: Guo, Yangyang, et al.
Published: (2026)
Split Unlearning
by: Yu, Guangsheng, et al.
Published: (2023)
by: Yu, Guangsheng, et al.
Published: (2023)
DFB: A Data-Free, Low-Budget, and High-Efficacy Clean-Label Backdoor Attack
by: Ma, Binhao, et al.
Published: (2023)
by: Ma, Binhao, et al.
Published: (2023)
DualBreach: Efficient Dual-Jailbreaking via Target-Driven Initialization and Multi-Target Optimization
by: Huang, Xinzhe, et al.
Published: (2025)
by: Huang, Xinzhe, et al.
Published: (2025)
Machine Unlearning in Large Language Models
by: Chen, Kongyang, et al.
Published: (2024)
by: Chen, Kongyang, et al.
Published: (2024)
The Real Menace of Cloning Attacks on SGX Applications
by: Wilde, Annika, et al.
Published: (2026)
by: Wilde, Annika, et al.
Published: (2026)
Pre-trained Encoder Inference: Revealing Upstream Encoders In Downstream Machine Learning Services
by: Fu, Shaopeng, et al.
Published: (2024)
by: Fu, Shaopeng, et al.
Published: (2024)
RedAgent: Red Teaming Large Language Models with Context-aware Autonomous Language Agent
by: Xu, Huiyu, et al.
Published: (2024)
by: Xu, Huiyu, et al.
Published: (2024)
Text-CRS: A Generalized Certified Robustness Framework against Textual Adversarial Attacks
by: Zhang, Xinyu, et al.
Published: (2023)
by: Zhang, Xinyu, et al.
Published: (2023)
Robust Watermarks Leak: Channel-Aware Feature Extraction Enables Adversarial Watermark Manipulation
by: Ba, Zhongjie, et al.
Published: (2025)
by: Ba, Zhongjie, et al.
Published: (2025)
MIRAGE: Misleading Retrieval-Augmented Generation via Black-box and Query-agnostic Poisoning Attacks
by: Chen, Tailun, et al.
Published: (2025)
by: Chen, Tailun, et al.
Published: (2025)
Textual Unlearning Gives a False Sense of Unlearning
by: Du, Jiacheng, et al.
Published: (2024)
by: Du, Jiacheng, et al.
Published: (2024)
Adversarial Machine Unlearning
by: Di, Zonglin, et al.
Published: (2024)
by: Di, Zonglin, et al.
Published: (2024)
Phoneme-Based Proactive Anti-Eavesdropping with Controlled Recording Privilege
by: Huang, Peng, et al.
Published: (2024)
by: Huang, Peng, et al.
Published: (2024)
ExpShield: Safeguarding Web Text from Unauthorized Crawling and LLM Exploitation
by: Liu, Ruixuan, et al.
Published: (2024)
by: Liu, Ruixuan, et al.
Published: (2024)
Certifiable Black-Box Attacks with Randomized Adversarial Examples: Breaking Defenses with Provable Confidence
by: Hong, Hanbin, et al.
Published: (2023)
by: Hong, Hanbin, et al.
Published: (2023)
Membership Inference Attacks Against Vision-Language Models
by: Hu, Yuke, et al.
Published: (2025)
by: Hu, Yuke, et al.
Published: (2025)
Defense against Poisoning Attacks under Shuffle-DP
by: Wang, Siyi, et al.
Published: (2026)
by: Wang, Siyi, et al.
Published: (2026)
Towards Classifying Benign And Malicious Packages Using Machine Learning
by: Nguyen, Thanh-Cong, et al.
Published: (2025)
by: Nguyen, Thanh-Cong, et al.
Published: (2025)
Unlearn and Burn: Adversarial Machine Unlearning Requests Destroy Model Accuracy
by: Huang, Yangsibo, et al.
Published: (2024)
by: Huang, Yangsibo, et al.
Published: (2024)
Really Unlearned? Verifying Machine Unlearning via Influential Sample Pairs
by: Xu, Heng, et al.
Published: (2024)
by: Xu, Heng, et al.
Published: (2024)
FDINet: Protecting against DNN Model Extraction via Feature Distortion Index
by: Yao, Hongwei, et al.
Published: (2023)
by: Yao, Hongwei, et al.
Published: (2023)
Explanation as a Watermark: Towards Harmless and Multi-bit Model Ownership Verification via Watermarking Feature Attribution
by: Shao, Shuo, et al.
Published: (2024)
by: Shao, Shuo, et al.
Published: (2024)
ALIF: Low-Cost Adversarial Audio Attacks on Black-Box Speech Platforms using Linguistic Features
by: Cheng, Peng, et al.
Published: (2024)
by: Cheng, Peng, et al.
Published: (2024)
CRFU: Compressive Representation Forgetting Against Privacy Leakage on Machine Unlearning
by: Wang, Weiqi, et al.
Published: (2025)
by: Wang, Weiqi, et al.
Published: (2025)
Similar Items
-
Learn What You Want to Unlearn: Unlearning Inversion Attacks against Machine Unlearning
by: Hu, Hongsheng, et al.
Published: (2024) -
MAJIC: Markovian Adaptive Jailbreaking via Iterative Composition of Diverse Innovative Strategies
by: Qi, Weiwei, et al.
Published: (2025) -
Untargeted Jailbreak Attack
by: Huang, Xinzhe, et al.
Published: (2025) -
ERASER: Machine Unlearning in MLaaS via an Inference Serving-Aware Approach
by: Hu, Yuke, et al.
Published: (2023) -
WMCopier: Forging Invisible Image Watermarks on Arbitrary Images
by: Dong, Ziping, et al.
Published: (2025)