Understanding Private Learning From Feature Perspective

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ding, Meng, Lei, Mingxi, Fu, Shaopeng, Wang, Shaowei, Wang, Di, Xu, Jinhui
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914167403839488
author Ding, Meng
Lei, Mingxi
Fu, Shaopeng
Wang, Shaowei
Wang, Di
Xu, Jinhui
author_facet Ding, Meng
Lei, Mingxi
Fu, Shaopeng
Wang, Shaowei
Wang, Di
Xu, Jinhui
contents Differentially private Stochastic Gradient Descent (DP-SGD) has become integral to privacy-preserving machine learning, ensuring robust privacy guarantees in sensitive domains. Despite notable empirical advances leveraging features from non-private, pre-trained models to enhance DP-SGD training, a theoretical understanding of feature dynamics in private learning remains underexplored. This paper presents the first theoretical framework to analyze private training through a feature learning perspective. Building on the multi-patch data structure from prior work, our analysis distinguishes between label-dependent feature signals and label-independent noise, a critical aspect overlooked by existing analyses in the DP community. Employing a two-layer CNN with polynomial ReLU activation, we theoretically characterize both feature signal learning and data noise memorization in private training via noisy gradient descent. Our findings reveal that (1) Effective private signal learning requires a higher signal-to-noise ratio (SNR) compared to non-private training, and (2) When data noise memorization occurs in non-private learning, it will also occur in private learning, leading to poor generalization despite small training loss. Our findings highlight the challenges of private learning and prove the benefit of feature enhancement to improve SNR. Experiments on synthetic and real-world datasets also validate our theoretical findings.
format Preprint
id arxiv_https___arxiv_org_abs_2511_18006
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Understanding Private Learning From Feature Perspective
Ding, Meng
Lei, Mingxi
Fu, Shaopeng
Wang, Shaowei
Wang, Di
Xu, Jinhui
Machine Learning
Differentially private Stochastic Gradient Descent (DP-SGD) has become integral to privacy-preserving machine learning, ensuring robust privacy guarantees in sensitive domains. Despite notable empirical advances leveraging features from non-private, pre-trained models to enhance DP-SGD training, a theoretical understanding of feature dynamics in private learning remains underexplored. This paper presents the first theoretical framework to analyze private training through a feature learning perspective. Building on the multi-patch data structure from prior work, our analysis distinguishes between label-dependent feature signals and label-independent noise, a critical aspect overlooked by existing analyses in the DP community. Employing a two-layer CNN with polynomial ReLU activation, we theoretically characterize both feature signal learning and data noise memorization in private training via noisy gradient descent. Our findings reveal that (1) Effective private signal learning requires a higher signal-to-noise ratio (SNR) compared to non-private training, and (2) When data noise memorization occurs in non-private learning, it will also occur in private learning, leading to poor generalization despite small training loss. Our findings highlight the challenges of private learning and prove the benefit of feature enhancement to improve SNR. Experiments on synthetic and real-world datasets also validate our theoretical findings.
title Understanding Private Learning From Feature Perspective
topic Machine Learning
url https://arxiv.org/abs/2511.18006