Towards Resilient Safety-driven Unlearning for Diffusion Models against Downstream Fine-tuning
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Boheng, Gu, Renjie, Wang, Junjie, Qi, Leyi, Li, Yiming, Wang, Run, Qin, Zhan, Zhang, Tianwei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
External Data Extraction Attacks against Retrieval-Augmented Large Language Models
by: He, Yu, et al.
Published: (2025)
by: He, Yu, et al.
Published: (2025)
DREAM: Scalable Red Teaming for Text-to-Image Generative Systems via Distribution Modeling
by: Li, Boheng, et al.
Published: (2025)
by: Li, Boheng, et al.
Published: (2025)
Taught Well Learned Ill: Towards Distillation-conditional Backdoor Attack
by: Chen, Yukun, et al.
Published: (2025)
by: Chen, Yukun, et al.
Published: (2025)
Towards Label-Only Membership Inference Attack against Pre-trained Large Language Models
by: He, Yu, et al.
Published: (2025)
by: He, Yu, et al.
Published: (2025)
Mitigating Fine-tuning based Jailbreak Attack with Backdoor Enhanced Safety Alignment
by: Wang, Jiongxiao, et al.
Published: (2024)
by: Wang, Jiongxiao, et al.
Published: (2024)
Towards Effective Prompt Stealing Attack against Text-to-Image Diffusion Models
by: Zhao, Shiqian, et al.
Published: (2025)
by: Zhao, Shiqian, et al.
Published: (2025)
Black-box Membership Inference Attacks against Fine-tuned Diffusion Models
by: Pang, Yan, et al.
Published: (2023)
by: Pang, Yan, et al.
Published: (2023)
ShadowCode: Towards (Automatic) External Prompt Injection Attack against Code LLMs
by: Yang, Yuchen, et al.
Published: (2024)
by: Yang, Yuchen, et al.
Published: (2024)
Nearest is Not Dearest: Towards Practical Defense against Quantization-conditioned Backdoor Attacks
by: Li, Boheng, et al.
Published: (2024)
by: Li, Boheng, et al.
Published: (2024)
BitHydra: Towards Bit-flip Inference Cost Attack against Large Language Models
by: Yan, Xiaobei, et al.
Published: (2025)
by: Yan, Xiaobei, et al.
Published: (2025)
Pharmacist: Safety Alignment Data Curation for Large Language Models against Harmful Fine-tuning
by: Liu, Guozhi, et al.
Published: (2025)
by: Liu, Guozhi, et al.
Published: (2025)
Robust-Wide: Robust Watermarking against Instruction-driven Image Editing
by: Hu, Runyi, et al.
Published: (2024)
by: Hu, Runyi, et al.
Published: (2024)
BURN: Backdoor Unlearning via Adversarial Boundary Analysis
by: Su, Yanghao, et al.
Published: (2025)
by: Su, Yanghao, et al.
Published: (2025)
Antidote: Post-fine-tuning Safety Alignment for Large Language Models against Harmful Fine-tuning
by: Huang, Tiansheng, et al.
Published: (2024)
by: Huang, Tiansheng, et al.
Published: (2024)
Learn What You Want to Unlearn: Unlearning Inversion Attacks against Machine Unlearning
by: Hu, Hongsheng, et al.
Published: (2024)
by: Hu, Hongsheng, et al.
Published: (2024)
Towards Action Hijacking of Large Language Model-based Agent
by: Zhang, Yuyang, et al.
Published: (2024)
by: Zhang, Yuyang, et al.
Published: (2024)
FIT-Print: Towards False-claim-resistant Model Ownership Verification via Targeted Fingerprint
by: Shao, Shuo, et al.
Published: (2025)
by: Shao, Shuo, et al.
Published: (2025)
Towards Lifecycle Unlearning Commitment Management: Measuring Sample-level Approximate Unlearning Completeness
by: Wang, Cheng-Long, et al.
Published: (2024)
by: Wang, Cheng-Long, et al.
Published: (2024)
EnchTable: Unified Safety Alignment Transfer in Fine-tuned Large Language Models
by: Wu, Jialin, et al.
Published: (2025)
by: Wu, Jialin, et al.
Published: (2025)
Shake to Leak: Fine-tuning Diffusion Models Can Amplify the Generative Privacy Risk
by: Li, Zhangheng, et al.
Published: (2024)
by: Li, Zhangheng, et al.
Published: (2024)
Label Inference Attacks against Federated Unlearning
by: Wang, Wei, et al.
Published: (2025)
by: Wang, Wei, et al.
Published: (2025)
Is Difficulty Calibration All We Need? Towards More Practical Membership Inference Attacks
by: He, Yu, et al.
Published: (2024)
by: He, Yu, et al.
Published: (2024)
Towards Identification and Intervention of Safety-Critical Parameters in Large Language Models
by: Qi, Weiwei, et al.
Published: (2026)
by: Qi, Weiwei, et al.
Published: (2026)
PoseGuard: Pose-Guided Generation with Safety Guardrails
by: Wang, Kongxin, et al.
Published: (2025)
by: Wang, Kongxin, et al.
Published: (2025)
Safety is Not Only About Refusal: Reasoning-Enhanced Fine-tuning for Interpretable LLM Safety
by: Zhang, Yuyou, et al.
Published: (2025)
by: Zhang, Yuyou, et al.
Published: (2025)
Towards Irreversible Machine Unlearning for Diffusion Models
by: Yuan, Xun, et al.
Published: (2025)
by: Yuan, Xun, et al.
Published: (2025)
Towards Lifecycle Unlearning Commitment Management: Measuring Sample-level Unlearning Completeness
by: Wang, Cheng-Long, et al.
Published: (2025)
by: Wang, Cheng-Long, et al.
Published: (2025)
Towards Physically Realizable Adversarial Attenuation Patch against SAR Object Detection
by: Zhang, Yiming, et al.
Published: (2026)
by: Zhang, Yiming, et al.
Published: (2026)
Cert-LAS: Toward Certified Model Ownership Verification for Text-to-Image Diffusion Models via Layer-Adaptive Smoothing
by: Qi, Leyi, et al.
Published: (2026)
by: Qi, Leyi, et al.
Published: (2026)
Split Unlearning
by: Yu, Guangsheng, et al.
Published: (2023)
by: Yu, Guangsheng, et al.
Published: (2023)
SafeAligner: Safety Alignment against Jailbreak Attacks via Response Disparity Guidance
by: Huang, Caishuang, et al.
Published: (2024)
by: Huang, Caishuang, et al.
Published: (2024)
StrTune: Data Dependence-based Code Slicing for Binary Similarity Detection with Fine-tuned Representation
by: He, Kaiyan, et al.
Published: (2024)
by: He, Kaiyan, et al.
Published: (2024)
SDD: Self-Degraded Defense against Malicious Fine-tuning
by: Chen, Zixuan, et al.
Published: (2025)
by: Chen, Zixuan, et al.
Published: (2025)
DP-SAPF: Saliency-Aware Parameter Fine-tuning of Public Models for Differentially Private Image Synthesis
by: Gong, Chen, et al.
Published: (2026)
by: Gong, Chen, et al.
Published: (2026)
SOFT: Selective Data Obfuscation for Protecting LLM Fine-tuning against Membership Inference Attacks
by: Zhang, Kaiyuan, et al.
Published: (2025)
by: Zhang, Kaiyuan, et al.
Published: (2025)
QUEEN: Query Unlearning against Model Extraction
by: Chen, Huajie, et al.
Published: (2024)
by: Chen, Huajie, et al.
Published: (2024)
Watermarking LLM-Generated Datasets in Downstream Tasks
by: Liu, Yugeng, et al.
Published: (2025)
by: Liu, Yugeng, et al.
Published: (2025)
Differentially Private Parameter-Efficient Fine-tuning for Large ASR Models
by: Liu, Hongbin, et al.
Published: (2024)
by: Liu, Hongbin, et al.
Published: (2024)
No Two Devils Alike: Unveiling Distinct Mechanisms of Fine-tuning Attacks
by: Leong, Chak Tou, et al.
Published: (2024)
by: Leong, Chak Tou, et al.
Published: (2024)
Moderator: Moderating Text-to-Image Diffusion Models through Fine-grained Context-based Policies
by: Wang, Peiran, et al.
Published: (2024)
by: Wang, Peiran, et al.
Published: (2024)
Similar Items
-
External Data Extraction Attacks against Retrieval-Augmented Large Language Models
by: He, Yu, et al.
Published: (2025) -
DREAM: Scalable Red Teaming for Text-to-Image Generative Systems via Distribution Modeling
by: Li, Boheng, et al.
Published: (2025) -
Taught Well Learned Ill: Towards Distillation-conditional Backdoor Attack
by: Chen, Yukun, et al.
Published: (2025) -
Towards Label-Only Membership Inference Attack against Pre-trained Large Language Models
by: He, Yu, et al.
Published: (2025) -
Mitigating Fine-tuning based Jailbreak Attack with Backdoor Enhanced Safety Alignment
by: Wang, Jiongxiao, et al.
Published: (2024)