Hollowed Net for On-Device Personalization of Text-to-Image Diffusion Models

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Cho, Wonguk, Choi, Seokeon, Das, Debasmit, Reisser, Matthias, Kim, Taesup, Yun, Sungrack, Porikli, Fatih
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866909376064782336
author Cho, Wonguk
Choi, Seokeon
Das, Debasmit
Reisser, Matthias
Kim, Taesup
Yun, Sungrack
Porikli, Fatih
author_facet Cho, Wonguk
Choi, Seokeon
Das, Debasmit
Reisser, Matthias
Kim, Taesup
Yun, Sungrack
Porikli, Fatih
contents Recent advancements in text-to-image diffusion models have enabled the personalization of these models to generate custom images from textual prompts. This paper presents an efficient LoRA-based personalization approach for on-device subject-driven generation, where pre-trained diffusion models are fine-tuned with user-specific data on resource-constrained devices. Our method, termed Hollowed Net, enhances memory efficiency during fine-tuning by modifying the architecture of a diffusion U-Net to temporarily remove a fraction of its deep layers, creating a hollowed structure. This approach directly addresses on-device memory constraints and substantially reduces GPU memory requirements for training, in contrast to previous methods that primarily focus on minimizing training steps and reducing the number of parameters to update. Additionally, the personalized Hollowed Net can be transferred back into the original U-Net, enabling inference without additional memory overhead. Quantitative and qualitative analyses demonstrate that our approach not only reduces training memory to levels as low as those required for inference but also maintains or improves personalization performance compared to existing methods.
format Preprint
id arxiv_https___arxiv_org_abs_2411_01179
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Hollowed Net for On-Device Personalization of Text-to-Image Diffusion Models
Cho, Wonguk
Choi, Seokeon
Das, Debasmit
Reisser, Matthias
Kim, Taesup
Yun, Sungrack
Porikli, Fatih
Computer Vision and Pattern Recognition
Artificial Intelligence
Graphics
Machine Learning
Recent advancements in text-to-image diffusion models have enabled the personalization of these models to generate custom images from textual prompts. This paper presents an efficient LoRA-based personalization approach for on-device subject-driven generation, where pre-trained diffusion models are fine-tuned with user-specific data on resource-constrained devices. Our method, termed Hollowed Net, enhances memory efficiency during fine-tuning by modifying the architecture of a diffusion U-Net to temporarily remove a fraction of its deep layers, creating a hollowed structure. This approach directly addresses on-device memory constraints and substantially reduces GPU memory requirements for training, in contrast to previous methods that primarily focus on minimizing training steps and reducing the number of parameters to update. Additionally, the personalized Hollowed Net can be transferred back into the original U-Net, enabling inference without additional memory overhead. Quantitative and qualitative analyses demonstrate that our approach not only reduces training memory to levels as low as those required for inference but also maintains or improves personalization performance compared to existing methods.
title Hollowed Net for On-Device Personalization of Text-to-Image Diffusion Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Graphics
Machine Learning
url https://arxiv.org/abs/2411.01179