Dataset Size Recovery from LoRA Weights

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Salama, Mohammad, Kahana, Jonathan, Horwitz, Eliahu, Hoshen, Yedid
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911935246630912
author Salama, Mohammad
Kahana, Jonathan
Horwitz, Eliahu
Hoshen, Yedid
author_facet Salama, Mohammad
Kahana, Jonathan
Horwitz, Eliahu
Hoshen, Yedid
contents Model inversion and membership inference attacks aim to reconstruct and verify the data which a model was trained on. However, they are not guaranteed to find all training samples as they do not know the size of the training set. In this paper, we introduce a new task: dataset size recovery, that aims to determine the number of samples used to train a model, directly from its weights. We then propose DSiRe, a method for recovering the number of images used to fine-tune a model, in the common case where fine-tuning uses LoRA. We discover that both the norm and the spectrum of the LoRA matrices are closely linked to the fine-tuning dataset size; we leverage this finding to propose a simple yet effective prediction algorithm. To evaluate dataset size recovery of LoRA weights, we develop and release a new benchmark, LoRA-WiSE, consisting of over 25000 weight snapshots from more than 2000 diverse LoRA fine-tuned models. Our best classifier can predict the number of fine-tuning images with a mean absolute error of 0.36 images, establishing the feasibility of this attack.
format Preprint
id arxiv_https___arxiv_org_abs_2406_19395
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Dataset Size Recovery from LoRA Weights
Salama, Mohammad
Kahana, Jonathan
Horwitz, Eliahu
Hoshen, Yedid
Computer Vision and Pattern Recognition
Model inversion and membership inference attacks aim to reconstruct and verify the data which a model was trained on. However, they are not guaranteed to find all training samples as they do not know the size of the training set. In this paper, we introduce a new task: dataset size recovery, that aims to determine the number of samples used to train a model, directly from its weights. We then propose DSiRe, a method for recovering the number of images used to fine-tune a model, in the common case where fine-tuning uses LoRA. We discover that both the norm and the spectrum of the LoRA matrices are closely linked to the fine-tuning dataset size; we leverage this finding to propose a simple yet effective prediction algorithm. To evaluate dataset size recovery of LoRA weights, we develop and release a new benchmark, LoRA-WiSE, consisting of over 25000 weight snapshots from more than 2000 diverse LoRA fine-tuned models. Our best classifier can predict the number of fine-tuning images with a mean absolute error of 0.36 images, establishing the feasibility of this attack.
title Dataset Size Recovery from LoRA Weights
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2406.19395