Delving into Differentially Private Transformer

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Ding, Youlong, Wu, Xueyang, Meng, Yining, Luo, Yonggang, Wang, Hao, Pan, Weike
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912001797652480
author Ding, Youlong
Wu, Xueyang
Meng, Yining
Luo, Yonggang
Wang, Hao
Pan, Weike
author_facet Ding, Youlong
Wu, Xueyang
Meng, Yining
Luo, Yonggang
Wang, Hao
Pan, Weike
contents Deep learning with differential privacy (DP) has garnered significant attention over the past years, leading to the development of numerous methods aimed at enhancing model accuracy and training efficiency. This paper delves into the problem of training Transformer models with differential privacy. Our treatment is modular: the logic is to `reduce' the problem of training DP Transformer to the more basic problem of training DP vanilla neural nets. The latter is better understood and amenable to many model-agnostic methods. Such `reduction' is done by first identifying the hardness unique to DP Transformer training: the attention distraction phenomenon and a lack of compatibility with existing techniques for efficient gradient clipping. To deal with these two issues, we propose the Re-Attention Mechanism and Phantom Clipping, respectively. We believe that our work not only casts new light on training DP Transformers but also promotes a modular treatment to advance research in the field of differentially private deep learning.
format Preprint
id arxiv_https___arxiv_org_abs_2405_18194
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Delving into Differentially Private Transformer
Ding, Youlong
Wu, Xueyang
Meng, Yining
Luo, Yonggang
Wang, Hao
Pan, Weike
Machine Learning
Cryptography and Security
Deep learning with differential privacy (DP) has garnered significant attention over the past years, leading to the development of numerous methods aimed at enhancing model accuracy and training efficiency. This paper delves into the problem of training Transformer models with differential privacy. Our treatment is modular: the logic is to `reduce' the problem of training DP Transformer to the more basic problem of training DP vanilla neural nets. The latter is better understood and amenable to many model-agnostic methods. Such `reduction' is done by first identifying the hardness unique to DP Transformer training: the attention distraction phenomenon and a lack of compatibility with existing techniques for efficient gradient clipping. To deal with these two issues, we propose the Re-Attention Mechanism and Phantom Clipping, respectively. We believe that our work not only casts new light on training DP Transformers but also promotes a modular treatment to advance research in the field of differentially private deep learning.
title Delving into Differentially Private Transformer
topic Machine Learning
Cryptography and Security
url https://arxiv.org/abs/2405.18194