Saved in:
Bibliographic Details
Main Authors: Sander, Tom, Yu, Yaodong, Sanjabi, Maziar, Durmus, Alain, Ma, Yi, Chaudhuri, Kamalika, Guo, Chuan
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2403.02506
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913567513509888
author Sander, Tom
Yu, Yaodong
Sanjabi, Maziar
Durmus, Alain
Ma, Yi
Chaudhuri, Kamalika
Guo, Chuan
author_facet Sander, Tom
Yu, Yaodong
Sanjabi, Maziar
Durmus, Alain
Ma, Yi
Chaudhuri, Kamalika
Guo, Chuan
contents Differentially private (DP) machine learning is considered the gold-standard solution for training a model from sensitive data while still preserving privacy. However, a major barrier to achieving this ideal is its sub-optimal privacy-accuracy trade-off, which is particularly visible in DP representation learning. Specifically, it has been shown that under modest privacy budgets, most models learn representations that are not significantly better than hand-crafted features. In this work, we show that effective DP representation learning can be done via image captioning and scaling up to internet-scale multimodal datasets. Through a series of engineering tricks, we successfully train a DP image captioner (DP-Cap) on a 233M subset of LAION-2B from scratch using a reasonable amount of computation, and obtaining unprecedented high-quality image features that can be used in a variety of downstream vision and vision-language tasks. For example, under a privacy budget of $\varepsilon=8$ for the LAION dataset, a linear classifier trained on top of learned DP-Cap features attains $65.8\%$ accuracy on ImageNet-1K, considerably improving the previous SOTA of $56.5\%$.
format Preprint
id arxiv_https___arxiv_org_abs_2403_02506
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Differentially Private Representation Learning via Image Captioning
Sander, Tom
Yu, Yaodong
Sanjabi, Maziar
Durmus, Alain
Ma, Yi
Chaudhuri, Kamalika
Guo, Chuan
Computer Vision and Pattern Recognition
Machine Learning
Differentially private (DP) machine learning is considered the gold-standard solution for training a model from sensitive data while still preserving privacy. However, a major barrier to achieving this ideal is its sub-optimal privacy-accuracy trade-off, which is particularly visible in DP representation learning. Specifically, it has been shown that under modest privacy budgets, most models learn representations that are not significantly better than hand-crafted features. In this work, we show that effective DP representation learning can be done via image captioning and scaling up to internet-scale multimodal datasets. Through a series of engineering tricks, we successfully train a DP image captioner (DP-Cap) on a 233M subset of LAION-2B from scratch using a reasonable amount of computation, and obtaining unprecedented high-quality image features that can be used in a variety of downstream vision and vision-language tasks. For example, under a privacy budget of $\varepsilon=8$ for the LAION dataset, a linear classifier trained on top of learned DP-Cap features attains $65.8\%$ accuracy on ImageNet-1K, considerably improving the previous SOTA of $56.5\%$.
title Differentially Private Representation Learning via Image Captioning
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2403.02506