DataShield: Safety-degrading Data Filtering for LLM Benign Instruction Fine-Tuning
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Junbo, Zhou, Qianli, Deng, Xinyang, Jiang, Wen, Pan, Jie, Zhu, Jinbiao |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Understanding and Mitigating Over-refusal for Large Language Models via Safety Representation
by: Zhang, Junbo, et al.
Published: (2025)
by: Zhang, Junbo, et al.
Published: (2025)
Cross-Task Defense: Instruction-Tuning LLMs for Content Safety
by: Fu, Yu, et al.
Published: (2024)
by: Fu, Yu, et al.
Published: (2024)
Sugar-Coated Poison: Benign Generation Unlocks LLM Jailbreaking
by: Wu, Yu-Hang, et al.
Published: (2025)
by: Wu, Yu-Hang, et al.
Published: (2025)
What is in Your Safe Data? Identifying Benign Data that Breaks Safety
by: He, Luxi, et al.
Published: (2024)
by: He, Luxi, et al.
Published: (2024)
Benign Fine-Tuning Breaks Safety Alignment in Audio LLMs
by: Roh, Jaechul, et al.
Published: (2026)
by: Roh, Jaechul, et al.
Published: (2026)
Generalization-Enhanced Code Vulnerability Detection via Multi-Task Instruction Fine-Tuning
by: Du, Xiaohu, et al.
Published: (2024)
by: Du, Xiaohu, et al.
Published: (2024)
SCOUT: A Defense Against Data Poisoning Attacks in Fine-Tuned Language Models
by: Afane, Mohamed, et al.
Published: (2025)
by: Afane, Mohamed, et al.
Published: (2025)
Safety is Not Only About Refusal: Reasoning-Enhanced Fine-tuning for Interpretable LLM Safety
by: Zhang, Yuyou, et al.
Published: (2025)
by: Zhang, Yuyou, et al.
Published: (2025)
Self-Jailbreaking: Language Models Can Reason Themselves Out of Safety Alignment After Benign Reasoning Training
by: Yong, Zheng-Xin, et al.
Published: (2025)
by: Yong, Zheng-Xin, et al.
Published: (2025)
Instructional Segment Embedding: Improving LLM Safety with Instruction Hierarchy
by: Wu, Tong, et al.
Published: (2024)
by: Wu, Tong, et al.
Published: (2024)
OneShield -- the Next Generation of LLM Guardrails
by: DeLuca, Chad, et al.
Published: (2025)
by: DeLuca, Chad, et al.
Published: (2025)
One Word at a Time: Incremental Completion Decomposition Breaks LLM Safety
by: Arif, Samee, et al.
Published: (2026)
by: Arif, Samee, et al.
Published: (2026)
Token-level Data Selection for Safe LLM Fine-tuning
by: Li, Yanping, et al.
Published: (2026)
by: Li, Yanping, et al.
Published: (2026)
Secure LLM Fine-Tuning via Safety-Aware Probing
by: Wu, Chengcan, et al.
Published: (2025)
by: Wu, Chengcan, et al.
Published: (2025)
AgentShield: Deception-based Compromise Detection for Tool-using LLM Agents
by: Rassul, Yassin H., et al.
Published: (2026)
by: Rassul, Yassin H., et al.
Published: (2026)
Defending Against Weight-Poisoning Backdoor Attacks for Parameter-Efficient Fine-Tuning
by: Zhao, Shuai, et al.
Published: (2024)
by: Zhao, Shuai, et al.
Published: (2024)
Layer-Aware Representation Filtering: Purifying Finetuning Data to Preserve LLM Safety Alignment
by: Li, Hao, et al.
Published: (2025)
by: Li, Hao, et al.
Published: (2025)
AdaSteer: Your Aligned LLM is Inherently an Adaptive Jailbreak Defender
by: Zhao, Weixiang, et al.
Published: (2025)
by: Zhao, Weixiang, et al.
Published: (2025)
The Art of (Mis)alignment: How Fine-Tuning Methods Effectively Misalign and Realign LLMs in Post-Training
by: Zhang, Rui, et al.
Published: (2026)
by: Zhang, Rui, et al.
Published: (2026)
TrojanPraise: Jailbreak LLMs via Benign Fine-Tuning
by: Xie, Zhixin, et al.
Published: (2026)
by: Xie, Zhixin, et al.
Published: (2026)
TuBA: Cross-Lingual Transferability of Backdoor Attacks in LLMs with Instruction Tuning
by: He, Xuanli, et al.
Published: (2024)
by: He, Xuanli, et al.
Published: (2024)
Rethinking Bottlenecks in Safety Fine-Tuning of Vision Language Models
by: Ding, Yi, et al.
Published: (2025)
by: Ding, Yi, et al.
Published: (2025)
A Comprehensive Survey in LLM(-Agent) Full Stack Safety: Data, Training and Deployment
by: Wang, Kun, et al.
Published: (2025)
by: Wang, Kun, et al.
Published: (2025)
Leaner Training, Lower Leakage: Revisiting Memorization in LLM Fine-Tuning with LoRA
by: Wang, Fei, et al.
Published: (2025)
by: Wang, Fei, et al.
Published: (2025)
Privacy-Preserving Parameter-Efficient Fine-Tuning for Large Language Model Services
by: Li, Yansong, et al.
Published: (2023)
by: Li, Yansong, et al.
Published: (2023)
Behavioral Canaries: Auditing Private Retrieved Context Usage in RL Fine-Tuning
by: Chen, Chaoran, et al.
Published: (2026)
by: Chen, Chaoran, et al.
Published: (2026)
Mitigating Fine-tuning based Jailbreak Attack with Backdoor Enhanced Safety Alignment
by: Wang, Jiongxiao, et al.
Published: (2024)
by: Wang, Jiongxiao, et al.
Published: (2024)
Why LLM Safety Guardrails Collapse After Fine-tuning: A Similarity Analysis Between Alignment and Fine-tuning Datasets
by: Hsiung, Lei, et al.
Published: (2025)
by: Hsiung, Lei, et al.
Published: (2025)
RewardDS: Privacy-Preserving Fine-Tuning for Large Language Models via Reward Driven Data Synthesis
by: Wang, Jianwei, et al.
Published: (2025)
by: Wang, Jianwei, et al.
Published: (2025)
EnchTable: Unified Safety Alignment Transfer in Fine-tuned Large Language Models
by: Wu, Jialin, et al.
Published: (2025)
by: Wu, Jialin, et al.
Published: (2025)
Emerging Safety Attack and Defense in Federated Instruction Tuning of Large Language Models
by: Ye, Rui, et al.
Published: (2024)
by: Ye, Rui, et al.
Published: (2024)
SafeTune: Mitigating Data Poisoning in LLM Fine-Tuning for RTL Code Generation
by: Rezakhani, Mahshid, et al.
Published: (2026)
by: Rezakhani, Mahshid, et al.
Published: (2026)
Root Defence Strategies: Ensuring Safety of LLM at the Decoding Level
by: Zeng, Xinyi, et al.
Published: (2024)
by: Zeng, Xinyi, et al.
Published: (2024)
CTRAP: Embedding Collapse Trap to Safeguard Large Language Models from Harmful Fine-Tuning
by: Yi, Biao, et al.
Published: (2025)
by: Yi, Biao, et al.
Published: (2025)
PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks
by: Shen, Guobin, et al.
Published: (2025)
by: Shen, Guobin, et al.
Published: (2025)
DAVE: A Policy-Enforcing LLM Spokesperson for Secure Multi-Document Data Sharing
by: Brinkhege, René, et al.
Published: (2026)
by: Brinkhege, René, et al.
Published: (2026)
Analysing Safety Risks in LLMs Fine-Tuned with Pseudo-Malicious Cyber Security Data
by: ElZemity, Adel, et al.
Published: (2025)
by: ElZemity, Adel, et al.
Published: (2025)
Safety Anchor: Defending Harmful Fine-tuning via Geometric Bottlenecks
by: Lu, Guoxin, et al.
Published: (2026)
by: Lu, Guoxin, et al.
Published: (2026)
Refining Input Guardrails: Enhancing LLM-as-a-Judge Efficiency Through Chain-of-Thought Fine-Tuning and Alignment
by: Rad, Melissa Kazemi, et al.
Published: (2025)
by: Rad, Melissa Kazemi, et al.
Published: (2025)
TWGuard: A Case Study of LLM Safety Guardrails for Localized Linguistic Contexts
by: Chu, Hua-Rong, et al.
Published: (2026)
by: Chu, Hua-Rong, et al.
Published: (2026)
Similar Items
-
Understanding and Mitigating Over-refusal for Large Language Models via Safety Representation
by: Zhang, Junbo, et al.
Published: (2025) -
Cross-Task Defense: Instruction-Tuning LLMs for Content Safety
by: Fu, Yu, et al.
Published: (2024) -
Sugar-Coated Poison: Benign Generation Unlocks LLM Jailbreaking
by: Wu, Yu-Hang, et al.
Published: (2025) -
What is in Your Safe Data? Identifying Benign Data that Breaks Safety
by: He, Luxi, et al.
Published: (2024) -
Benign Fine-Tuning Breaks Safety Alignment in Audio LLMs
by: Roh, Jaechul, et al.
Published: (2026)