A New Linear Scaling Rule for Private Adaptive Hyperparameter Optimization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Panda, Ashwinee, Tang, Xinyu, Mahloujifar, Saeed, Sehwag, Vikash, Mittal, Prateek
Format: Preprint
Published: 2022
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916235810177024
author Panda, Ashwinee
Tang, Xinyu
Mahloujifar, Saeed
Sehwag, Vikash
Mittal, Prateek
author_facet Panda, Ashwinee
Tang, Xinyu
Mahloujifar, Saeed
Sehwag, Vikash
Mittal, Prateek
contents An open problem in differentially private deep learning is hyperparameter optimization (HPO). DP-SGD introduces new hyperparameters and complicates existing ones, forcing researchers to painstakingly tune hyperparameters with hundreds of trials, which in turn makes it impossible to account for the privacy cost of HPO without destroying the utility. We propose an adaptive HPO method that uses cheap trials (in terms of privacy cost and runtime) to estimate optimal hyperparameters and scales them up. We obtain state-of-the-art performance on 22 benchmark tasks, across computer vision and natural language processing, across pretraining and finetuning, across architectures and a wide range of $\varepsilon \in [0.01,8.0]$, all while accounting for the privacy cost of HPO.
format Preprint
id arxiv_https___arxiv_org_abs_2212_04486
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle A New Linear Scaling Rule for Private Adaptive Hyperparameter Optimization
Panda, Ashwinee
Tang, Xinyu
Mahloujifar, Saeed
Sehwag, Vikash
Mittal, Prateek
Machine Learning
Artificial Intelligence
Cryptography and Security
An open problem in differentially private deep learning is hyperparameter optimization (HPO). DP-SGD introduces new hyperparameters and complicates existing ones, forcing researchers to painstakingly tune hyperparameters with hundreds of trials, which in turn makes it impossible to account for the privacy cost of HPO without destroying the utility. We propose an adaptive HPO method that uses cheap trials (in terms of privacy cost and runtime) to estimate optimal hyperparameters and scales them up. We obtain state-of-the-art performance on 22 benchmark tasks, across computer vision and natural language processing, across pretraining and finetuning, across architectures and a wide range of $\varepsilon \in [0.01,8.0]$, all while accounting for the privacy cost of HPO.
title A New Linear Scaling Rule for Private Adaptive Hyperparameter Optimization
topic Machine Learning
Artificial Intelligence
Cryptography and Security
url https://arxiv.org/abs/2212.04486