Differentially Private Subspace Fine-Tuning for Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zheng, Lele, Wang, Xiang, Zhang, Tao, Cao, Yang, Cheng, Ke, Shen, Yulong
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912827925594112
author Zheng, Lele
Wang, Xiang
Zhang, Tao
Cao, Yang
Cheng, Ke
Shen, Yulong
author_facet Zheng, Lele
Wang, Xiang
Zhang, Tao
Cao, Yang
Cheng, Ke
Shen, Yulong
contents Fine-tuning large language models on downstream tasks is crucial for realizing their cross-domain potential but often relies on sensitive data, raising privacy concerns. Differential privacy (DP) offers rigorous privacy guarantees and has been widely adopted in fine-tuning; however, naively injecting noise across the high-dimensional parameter space creates perturbations with large norms, degrading performance and destabilizing training. To address this issue, we propose DP-SFT, a two-stage subspace fine-tuning method that substantially reduces noise magnitude while preserving formal DP guarantees. Our intuition is that, during fine-tuning, significant parameter updates lie within a low-dimensional, task-specific subspace, while other directions change minimally. Hence, we only inject DP noise into this subspace to protect privacy without perturbing irrelevant parameters. In phase one, we identify the subspace by analyzing principal gradient directions to capture task-specific update signals. In phase two, we project full gradients onto this subspace, add DP noise, and map the perturbed gradients back to the original parameter space for model updates, markedly lowering noise impact. Experiments on multiple datasets demonstrate that DP-SFT enhances accuracy and stability under rigorous DP constraints, accelerates convergence, and achieves substantial gains over DP fine-tuning baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2601_11113
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Differentially Private Subspace Fine-Tuning for Large Language Models
Zheng, Lele
Wang, Xiang
Zhang, Tao
Cao, Yang
Cheng, Ke
Shen, Yulong
Machine Learning
Cryptography and Security
Fine-tuning large language models on downstream tasks is crucial for realizing their cross-domain potential but often relies on sensitive data, raising privacy concerns. Differential privacy (DP) offers rigorous privacy guarantees and has been widely adopted in fine-tuning; however, naively injecting noise across the high-dimensional parameter space creates perturbations with large norms, degrading performance and destabilizing training. To address this issue, we propose DP-SFT, a two-stage subspace fine-tuning method that substantially reduces noise magnitude while preserving formal DP guarantees. Our intuition is that, during fine-tuning, significant parameter updates lie within a low-dimensional, task-specific subspace, while other directions change minimally. Hence, we only inject DP noise into this subspace to protect privacy without perturbing irrelevant parameters. In phase one, we identify the subspace by analyzing principal gradient directions to capture task-specific update signals. In phase two, we project full gradients onto this subspace, add DP noise, and map the perturbed gradients back to the original parameter space for model updates, markedly lowering noise impact. Experiments on multiple datasets demonstrate that DP-SFT enhances accuracy and stability under rigorous DP constraints, accelerates convergence, and achieves substantial gains over DP fine-tuning baselines.
title Differentially Private Subspace Fine-Tuning for Large Language Models
topic Machine Learning
Cryptography and Security
url https://arxiv.org/abs/2601.11113