Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets
Fuente:
arXiv
Saved in:
| Main Authors: | Lu, Ning, Liu, Shengcai, Wu, Jiahao, Chen, Weiyu, Zhang, Zhirui, Ong, Yew-Soon, Wang, Qi, Tang, Ke |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Less is More: Understanding Word-level Textual Adversarial Attack via n-gram Frequency Descend
by: Lu, Ning, et al.
Published: (2023)
by: Lu, Ning, et al.
Published: (2023)
Benign Fine-Tuning Breaks Safety Alignment in Audio LLMs
by: Roh, Jaechul, et al.
Published: (2026)
by: Roh, Jaechul, et al.
Published: (2026)
Evolving Jailbreaks: Automated Multi-Objective Long-Tail Attacks on Large Language Models
by: Hong, Wenjing, et al.
Published: (2026)
by: Hong, Wenjing, et al.
Published: (2026)
Backdoor Graph Condensation
by: Wu, Jiahao, et al.
Published: (2024)
by: Wu, Jiahao, et al.
Published: (2024)
FedRLHF: A Convergence-Guaranteed Federated Framework for Privacy-Preserving and Personalized RLHF
by: Fan, Flint Xiaofeng, et al.
Published: (2024)
by: Fan, Flint Xiaofeng, et al.
Published: (2024)
RefusalGuard: Geometry-Preserving Fine-Tuning for Safety in LLMs
by: Asif, Sadia, et al.
Published: (2026)
by: Asif, Sadia, et al.
Published: (2026)
Analysing Safety Risks in LLMs Fine-Tuned with Pseudo-Malicious Cyber Security Data
by: ElZemity, Adel, et al.
Published: (2025)
by: ElZemity, Adel, et al.
Published: (2025)
Defending MoE LLMs against Harmful Fine-Tuning via Safety Routing Alignment
by: Kim, Jaehan, et al.
Published: (2025)
by: Kim, Jaehan, et al.
Published: (2025)
SafeTune: Mitigating Data Poisoning in LLM Fine-Tuning for RTL Code Generation
by: Rezakhani, Mahshid, et al.
Published: (2026)
by: Rezakhani, Mahshid, et al.
Published: (2026)
Guard-GBDT: Efficient Privacy-Preserving Approximated GBDT Training on Vertical Dataset
by: Song, Anxiao, et al.
Published: (2025)
by: Song, Anxiao, et al.
Published: (2025)
Fine-Tuning, Quantization, and LLMs: Navigating Unintended Outcomes
by: Kumar, Divyanshu, et al.
Published: (2024)
by: Kumar, Divyanshu, et al.
Published: (2024)
When FinTech Meets Privacy: Securing Financial LLMs with Differential Private Fine-Tuning
by: Zhu, Sichen, et al.
Published: (2025)
by: Zhu, Sichen, et al.
Published: (2025)
Privacy-Preserving Parameter-Efficient Fine-Tuning for Large Language Model Services
by: Li, Yansong, et al.
Published: (2023)
by: Li, Yansong, et al.
Published: (2023)
Cross-Task Defense: Instruction-Tuning LLMs for Content Safety
by: Fu, Yu, et al.
Published: (2024)
by: Fu, Yu, et al.
Published: (2024)
SPARD: Defending Harmful Fine-Tuning Attack via Safety Projection with Relevance-Diversity Data Selection
by: Chen, Shuhao, et al.
Published: (2026)
by: Chen, Shuhao, et al.
Published: (2026)
TrojanPraise: Jailbreak LLMs via Benign Fine-Tuning
by: Xie, Zhixin, et al.
Published: (2026)
by: Xie, Zhixin, et al.
Published: (2026)
Persistent Backdoor Attacks under Continual Fine-Tuning of LLMs
by: Cui, Jing, et al.
Published: (2025)
by: Cui, Jing, et al.
Published: (2025)
SafeCOMM: A Study on Safety Degradation in Fine-Tuned Telecom Large Language Models
by: Djuhera, Aladin, et al.
Published: (2025)
by: Djuhera, Aladin, et al.
Published: (2025)
Self-Mined Hardness for Safety Fine-Tuning
by: Gupta, Prakhar, et al.
Published: (2026)
by: Gupta, Prakhar, et al.
Published: (2026)
Scam Shield: Multi-Model Voting and Fine-Tuned LLMs Against Adversarial Attacks
by: Chang, Chen-Wei, et al.
Published: (2025)
by: Chang, Chen-Wei, et al.
Published: (2025)
GradSafe: Detecting Jailbreak Prompts for LLMs via Safety-Critical Gradient Analysis
by: Xie, Yueqi, et al.
Published: (2024)
by: Xie, Yueqi, et al.
Published: (2024)
Layer-Aware Representation Filtering: Purifying Finetuning Data to Preserve LLM Safety Alignment
by: Li, Hao, et al.
Published: (2025)
by: Li, Hao, et al.
Published: (2025)
Do Fine-Tuned LLMs Understand Vulnerabilities? An Investigation into the Semantic Trap
by: Huang, Feiyang, et al.
Published: (2026)
by: Huang, Feiyang, et al.
Published: (2026)
Differentiation-Based Extraction of Proprietary Data from Fine-Tuned LLMs
by: Li, Zongjie, et al.
Published: (2025)
by: Li, Zongjie, et al.
Published: (2025)
SafeToolBench: Pioneering a Prospective Benchmark to Evaluating Tool Utilization Safety in LLMs
by: Xia, Hongfei, et al.
Published: (2025)
by: Xia, Hongfei, et al.
Published: (2025)
RouteMark: A Fingerprint for Intellectual Property Attribution in Routing-based Model Merging
by: He, Xin, et al.
Published: (2025)
by: He, Xin, et al.
Published: (2025)
VulReaD: Knowledge-Graph-guided Software Vulnerability Reasoning and Detection
by: Mukhtar, Samal, et al.
Published: (2026)
by: Mukhtar, Samal, et al.
Published: (2026)
Data Provenance Auditing of Fine-Tuned Large Language Models with a Text-Preserving Technique
by: Li, Yanming, et al.
Published: (2025)
by: Li, Yanming, et al.
Published: (2025)
Fine-Tuning Foundation Models with Federated Learning for Privacy Preserving Medical Time Series Forecasting
by: Ali, Mahad, et al.
Published: (2025)
by: Ali, Mahad, et al.
Published: (2025)
TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering
by: Hossain, Saad, et al.
Published: (2026)
by: Hossain, Saad, et al.
Published: (2026)
Verifiable, Efficient and Confidentiality-Preserving Graph Search with Transparency
by: Wang, Qiuhao, et al.
Published: (2025)
by: Wang, Qiuhao, et al.
Published: (2025)
BugWhisperer: Fine-Tuning LLMs for SoC Hardware Vulnerability Detection
by: Tarek, Shams, et al.
Published: (2025)
by: Tarek, Shams, et al.
Published: (2025)
Fine-Tuning LLMs for Code Mutation: A New Era of Cyber Threats
by: Setak, Mohammad, et al.
Published: (2024)
by: Setak, Mohammad, et al.
Published: (2024)
Differentially Private Subspace Fine-Tuning for Large Language Models
by: Zheng, Lele, et al.
Published: (2026)
by: Zheng, Lele, et al.
Published: (2026)
ShapeMark: Robust and Diversity-Preserving Watermarking for Diffusion Models
by: Qian, Yuqi, et al.
Published: (2026)
by: Qian, Yuqi, et al.
Published: (2026)
Privacy-Preserving Logistic Regression Training on Large Datasets
by: Chiang, John
Published: (2024)
by: Chiang, John
Published: (2024)
SafeBench: A Safety Evaluation Framework for Multimodal Large Language Models
by: Ying, Zonghao, et al.
Published: (2024)
by: Ying, Zonghao, et al.
Published: (2024)
PrivTuner with Homomorphic Encryption and LoRA: A P3EFT Scheme for Privacy-Preserving Parameter-Efficient Fine-Tuning of AI Foundation Models
by: Li, Yang, et al.
Published: (2024)
by: Li, Yang, et al.
Published: (2024)
The Hidden Costs of Domain Fine-Tuning: Pii-Bearing Data Degrades Safety and Increases Leakage
by: Choudhari, Jayesh, et al.
Published: (2026)
by: Choudhari, Jayesh, et al.
Published: (2026)
A Systematic Evaluation of Parameter-Efficient Fine-Tuning Methods for the Security of Code LLMs
by: Lee, Kiho, et al.
Published: (2025)
by: Lee, Kiho, et al.
Published: (2025)
Similar Items
-
Less is More: Understanding Word-level Textual Adversarial Attack via n-gram Frequency Descend
by: Lu, Ning, et al.
Published: (2023) -
Benign Fine-Tuning Breaks Safety Alignment in Audio LLMs
by: Roh, Jaechul, et al.
Published: (2026) -
Evolving Jailbreaks: Automated Multi-Objective Long-Tail Attacks on Large Language Models
by: Hong, Wenjing, et al.
Published: (2026) -
Backdoor Graph Condensation
by: Wu, Jiahao, et al.
Published: (2024) -
FedRLHF: A Convergence-Guaranteed Federated Framework for Privacy-Preserving and Personalized RLHF
by: Fan, Flint Xiaofeng, et al.
Published: (2024)