DEFT: Decompositional Efficient Fine-Tuning for Text-to-Image Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kumar, Komal, Anwer, Rao Muhammad, Khan, Fahad Shahbaz, Khan, Salman, Laptev, Ivan, Cholakkal, Hisham
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912609426472960
author Kumar, Komal
Anwer, Rao Muhammad
Khan, Fahad Shahbaz
Khan, Salman
Laptev, Ivan
Cholakkal, Hisham
author_facet Kumar, Komal
Anwer, Rao Muhammad
Khan, Fahad Shahbaz
Khan, Salman
Laptev, Ivan
Cholakkal, Hisham
contents Efficient fine-tuning of pre-trained Text-to-Image (T2I) models involves adjusting the model to suit a particular task or dataset while minimizing computational resources and limiting the number of trainable parameters. However, it often faces challenges in striking a trade-off between aligning with the target distribution: learning a novel concept from a limited image for personalization and retaining the instruction ability needed for unifying multiple tasks, all while maintaining editability (aligning with a variety of prompts or in-context generation). In this work, we introduce DEFT, Decompositional Efficient Fine-Tuning, an efficient fine-tuning framework that adapts a pre-trained weight matrix by decomposing its update into two components with two trainable matrices: (1) a projection onto the complement of a low-rank subspace spanned by a low-rank matrix, and (2) a low-rank update. The single trainable low-rank matrix defines the subspace, while the other trainable low-rank matrix enables flexible parameter adaptation within that subspace. We conducted extensive experiments on the Dreambooth and Dreambench Plus datasets for personalization, the InsDet dataset for object and scene adaptation, and the VisualCloze dataset for a universal image generation framework through visual in-context learning with both Stable Diffusion and a unified model. Our results demonstrated state-of-the-art performance, highlighting the emergent properties of efficient fine-tuning. Our code is available on \href{https://github.com/MAXNORM8650/DEFT}{DEFTBase}.
format Preprint
id arxiv_https___arxiv_org_abs_2509_22793
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DEFT: Decompositional Efficient Fine-Tuning for Text-to-Image Models
Kumar, Komal
Anwer, Rao Muhammad
Khan, Fahad Shahbaz
Khan, Salman
Laptev, Ivan
Cholakkal, Hisham
Computer Vision and Pattern Recognition
Efficient fine-tuning of pre-trained Text-to-Image (T2I) models involves adjusting the model to suit a particular task or dataset while minimizing computational resources and limiting the number of trainable parameters. However, it often faces challenges in striking a trade-off between aligning with the target distribution: learning a novel concept from a limited image for personalization and retaining the instruction ability needed for unifying multiple tasks, all while maintaining editability (aligning with a variety of prompts or in-context generation). In this work, we introduce DEFT, Decompositional Efficient Fine-Tuning, an efficient fine-tuning framework that adapts a pre-trained weight matrix by decomposing its update into two components with two trainable matrices: (1) a projection onto the complement of a low-rank subspace spanned by a low-rank matrix, and (2) a low-rank update. The single trainable low-rank matrix defines the subspace, while the other trainable low-rank matrix enables flexible parameter adaptation within that subspace. We conducted extensive experiments on the Dreambooth and Dreambench Plus datasets for personalization, the InsDet dataset for object and scene adaptation, and the VisualCloze dataset for a universal image generation framework through visual in-context learning with both Stable Diffusion and a unified model. Our results demonstrated state-of-the-art performance, highlighting the emergent properties of efficient fine-tuning. Our code is available on \href{https://github.com/MAXNORM8650/DEFT}{DEFTBase}.
title DEFT: Decompositional Efficient Fine-Tuning for Text-to-Image Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.22793