TARA: Token-Aware LoRA for Composable Personalization in Diffusion Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Peng, Yuqi, Zheng, Lingtao, Yang, Yufeng, Huang, Yi, Yan, Mingfu, Liu, Jianzhuang, Chen, Shifeng
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916892999942144
author Peng, Yuqi
Zheng, Lingtao
Yang, Yufeng
Huang, Yi
Yan, Mingfu
Liu, Jianzhuang
Chen, Shifeng
author_facet Peng, Yuqi
Zheng, Lingtao
Yang, Yufeng
Huang, Yi
Yan, Mingfu
Liu, Jianzhuang
Chen, Shifeng
contents Personalized text-to-image generation aims to synthesize novel images of a specific subject or style using only a few reference images. Recent methods based on Low-Rank Adaptation (LoRA) enable efficient single-concept customization by injecting lightweight, concept-specific adapters into pre-trained diffusion models. However, combining multiple LoRA modules for multi-concept generation often leads to identity missing and visual feature leakage. In this work, we identify two key issues behind these failures: (1) token-wise interference among different LoRA modules, and (2) spatial misalignment between the attention map of a rare token and its corresponding concept-specific region. To address these issues, we propose Token-Aware LoRA (TARA), which introduces a token mask to explicitly constrain each module to focus on its associated rare token to avoid interference, and a training objective that encourages the spatial attention of a rare token to align with its concept region. Our method enables training-free multi-concept composition by directly injecting multiple independently trained TARA modules at inference time. Experimental results demonstrate that TARA enables efficient multi-concept inference and effectively preserving the visual identity of each concept by avoiding mutual interference between LoRA modules. The code and models are available at https://github.com/YuqiPeng77/TARA.
format Preprint
id arxiv_https___arxiv_org_abs_2508_08812
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TARA: Token-Aware LoRA for Composable Personalization in Diffusion Models
Peng, Yuqi
Zheng, Lingtao
Yang, Yufeng
Huang, Yi
Yan, Mingfu
Liu, Jianzhuang
Chen, Shifeng
Computer Vision and Pattern Recognition
Personalized text-to-image generation aims to synthesize novel images of a specific subject or style using only a few reference images. Recent methods based on Low-Rank Adaptation (LoRA) enable efficient single-concept customization by injecting lightweight, concept-specific adapters into pre-trained diffusion models. However, combining multiple LoRA modules for multi-concept generation often leads to identity missing and visual feature leakage. In this work, we identify two key issues behind these failures: (1) token-wise interference among different LoRA modules, and (2) spatial misalignment between the attention map of a rare token and its corresponding concept-specific region. To address these issues, we propose Token-Aware LoRA (TARA), which introduces a token mask to explicitly constrain each module to focus on its associated rare token to avoid interference, and a training objective that encourages the spatial attention of a rare token to align with its concept region. Our method enables training-free multi-concept composition by directly injecting multiple independently trained TARA modules at inference time. Experimental results demonstrate that TARA enables efficient multi-concept inference and effectively preserving the visual identity of each concept by avoiding mutual interference between LoRA modules. The code and models are available at https://github.com/YuqiPeng77/TARA.
title TARA: Token-Aware LoRA for Composable Personalization in Diffusion Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.08812