DAGER: Exact Gradient Inversion for Large Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Petrov, Ivo, Dimitrov, Dimitar I., Baader, Maximilian, Müller, Mark Niklas, Vechev, Martin
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909387436589056
author Petrov, Ivo
Dimitrov, Dimitar I.
Baader, Maximilian
Müller, Mark Niklas
Vechev, Martin
author_facet Petrov, Ivo
Dimitrov, Dimitar I.
Baader, Maximilian
Müller, Mark Niklas
Vechev, Martin
contents Federated learning works by aggregating locally computed gradients from multiple clients, thus enabling collaborative training without sharing private client data. However, prior work has shown that the data can actually be recovered by the server using so-called gradient inversion attacks. While these attacks perform well when applied on images, they are limited in the text domain and only permit approximate reconstruction of small batches and short input sequences. In this work, we propose DAGER, the first algorithm to recover whole batches of input text exactly. DAGER leverages the low-rank structure of self-attention layer gradients and the discrete nature of token embeddings to efficiently check if a given token sequence is part of the client data. We use this check to exactly recover full batches in the honest-but-curious setting without any prior on the data for both encoder- and decoder-based architectures using exhaustive heuristic search and a greedy approach, respectively. We provide an efficient GPU implementation of DAGER and show experimentally that it recovers full batches of size up to 128 on large language models (LLMs), beating prior attacks in speed (20x at same batch size), scalability (10x larger batches), and reconstruction quality (ROUGE-1/2 > 0.99).
format Preprint
id arxiv_https___arxiv_org_abs_2405_15586
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle DAGER: Exact Gradient Inversion for Large Language Models
Petrov, Ivo
Dimitrov, Dimitar I.
Baader, Maximilian
Müller, Mark Niklas
Vechev, Martin
Machine Learning
Distributed, Parallel, and Cluster Computing
I.2.7; I.2.11
Federated learning works by aggregating locally computed gradients from multiple clients, thus enabling collaborative training without sharing private client data. However, prior work has shown that the data can actually be recovered by the server using so-called gradient inversion attacks. While these attacks perform well when applied on images, they are limited in the text domain and only permit approximate reconstruction of small batches and short input sequences. In this work, we propose DAGER, the first algorithm to recover whole batches of input text exactly. DAGER leverages the low-rank structure of self-attention layer gradients and the discrete nature of token embeddings to efficiently check if a given token sequence is part of the client data. We use this check to exactly recover full batches in the honest-but-curious setting without any prior on the data for both encoder- and decoder-based architectures using exhaustive heuristic search and a greedy approach, respectively. We provide an efficient GPU implementation of DAGER and show experimentally that it recovers full batches of size up to 128 on large language models (LLMs), beating prior attacks in speed (20x at same batch size), scalability (10x larger batches), and reconstruction quality (ROUGE-1/2 > 0.99).
title DAGER: Exact Gradient Inversion for Large Language Models
topic Machine Learning
Distributed, Parallel, and Cluster Computing
I.2.7; I.2.11
url https://arxiv.org/abs/2405.15586