DMin: Scalable Training Data Influence Estimation for Diffusion Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Lin, Huawei, Lao, Yingjie, Zhao, Weijie
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914459574861824
author Lin, Huawei
Lao, Yingjie
Zhao, Weijie
author_facet Lin, Huawei
Lao, Yingjie
Zhao, Weijie
contents Identifying the training data samples that most influence a generated image is a critical task in understanding diffusion models (DMs), yet existing influence estimation methods are constrained to small-scale or LoRA-tuned models due to computational limitations. To address this challenge, we propose DMin (Diffusion Model influence), a scalable framework for estimating the influence of each training data sample on a given generated image. To the best of our knowledge, it is the first method capable of influence estimation for DMs with billions of parameters. Leveraging efficient gradient compression, DMin reduces storage requirements from hundreds of TBs to mere MBs or even KBs, and retrieves the top-k most influential training samples in under 1 second, all while maintaining performance. Our empirical results demonstrate DMin is both effective in identifying influential training samples and efficient in terms of computational and storage requirements.
format Preprint
id arxiv_https___arxiv_org_abs_2412_08637
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle DMin: Scalable Training Data Influence Estimation for Diffusion Models
Lin, Huawei
Lao, Yingjie
Zhao, Weijie
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Identifying the training data samples that most influence a generated image is a critical task in understanding diffusion models (DMs), yet existing influence estimation methods are constrained to small-scale or LoRA-tuned models due to computational limitations. To address this challenge, we propose DMin (Diffusion Model influence), a scalable framework for estimating the influence of each training data sample on a given generated image. To the best of our knowledge, it is the first method capable of influence estimation for DMs with billions of parameters. Leveraging efficient gradient compression, DMin reduces storage requirements from hundreds of TBs to mere MBs or even KBs, and retrieves the top-k most influential training samples in under 1 second, all while maintaining performance. Our empirical results demonstrate DMin is both effective in identifying influential training samples and efficient in terms of computational and storage requirements.
title DMin: Scalable Training Data Influence Estimation for Diffusion Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2412.08637