Privacy-Preserving Transformers: SwiftKey's Differential Privacy Implementation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866916728375607296 |
|---|---|
| author | Abouelenin, Abdelrahman Abdelrehim, Mohamed Fahim, Raffy Hendy, Amr Afify, Mohamed |
| author_facet | Abouelenin, Abdelrahman Abdelrehim, Mohamed Fahim, Raffy Hendy, Amr Afify, Mohamed |
| contents | In this paper we train a transformer using differential privacy (DP) for language modeling in SwiftKey. We run multiple experiments to balance the trade-off between the model size, run-time speed and accuracy. We show that we get small and consistent gains in the next-word-prediction and accuracy with graceful increase in memory and speed compared to the production GRU. This is obtained by scaling down a GPT2 architecture to fit the required size and a two stage training process that builds a seed model on general data and DP finetunes it on typing data. The transformer is integrated using ONNX offering both flexibility and efficiency. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_05648 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Privacy-Preserving Transformers: SwiftKey's Differential Privacy Implementation Abouelenin, Abdelrahman Abdelrehim, Mohamed Fahim, Raffy Hendy, Amr Afify, Mohamed Computation and Language Cryptography and Security Machine Learning In this paper we train a transformer using differential privacy (DP) for language modeling in SwiftKey. We run multiple experiments to balance the trade-off between the model size, run-time speed and accuracy. We show that we get small and consistent gains in the next-word-prediction and accuracy with graceful increase in memory and speed compared to the production GRU. This is obtained by scaling down a GPT2 architecture to fit the required size and a two stage training process that builds a seed model on general data and DP finetunes it on typing data. The transformer is integrated using ONNX offering both flexibility and efficiency. |
| title | Privacy-Preserving Transformers: SwiftKey's Differential Privacy Implementation |
| topic | Computation and Language Cryptography and Security Machine Learning |
| url | https://arxiv.org/abs/2505.05648 |