MultLFG: Training-free Multi-LoRA composition using Frequency-domain Guidance

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Roy, Aniket, Suin, Maitreya, Shah, Ketul, Chellappa, Rama
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916760305795072
author Roy, Aniket
Suin, Maitreya
Shah, Ketul
Chellappa, Rama
author_facet Roy, Aniket
Suin, Maitreya
Shah, Ketul
Chellappa, Rama
contents Low-Rank Adaptation (LoRA) has gained prominence as a computationally efficient method for fine-tuning generative models, enabling distinct visual concept synthesis with minimal overhead. However, current methods struggle to effectively merge multiple LoRA adapters without training, particularly in complex compositions involving diverse visual elements. We introduce MultLFG, a novel framework for training-free multi-LoRA composition that utilizes frequency-domain guidance to achieve adaptive fusion of multiple LoRAs. Unlike existing methods that uniformly aggregate concept-specific LoRAs, MultLFG employs a timestep and frequency subband adaptive fusion strategy, selectively activating relevant LoRAs based on content relevance at specific timesteps and frequency bands. This frequency-sensitive guidance not only improves spatial coherence but also provides finer control over multi-LoRA composition, leading to more accurate and consistent results. Experimental evaluations on the ComposLoRA benchmark reveal that MultLFG substantially enhances compositional fidelity and image quality across various styles and concept sets, outperforming state-of-the-art baselines in multi-concept generation tasks. Code will be released.
format Preprint
id arxiv_https___arxiv_org_abs_2505_20525
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MultLFG: Training-free Multi-LoRA composition using Frequency-domain Guidance
Roy, Aniket
Suin, Maitreya
Shah, Ketul
Chellappa, Rama
Computer Vision and Pattern Recognition
Low-Rank Adaptation (LoRA) has gained prominence as a computationally efficient method for fine-tuning generative models, enabling distinct visual concept synthesis with minimal overhead. However, current methods struggle to effectively merge multiple LoRA adapters without training, particularly in complex compositions involving diverse visual elements. We introduce MultLFG, a novel framework for training-free multi-LoRA composition that utilizes frequency-domain guidance to achieve adaptive fusion of multiple LoRAs. Unlike existing methods that uniformly aggregate concept-specific LoRAs, MultLFG employs a timestep and frequency subband adaptive fusion strategy, selectively activating relevant LoRAs based on content relevance at specific timesteps and frequency bands. This frequency-sensitive guidance not only improves spatial coherence but also provides finer control over multi-LoRA composition, leading to more accurate and consistent results. Experimental evaluations on the ComposLoRA benchmark reveal that MultLFG substantially enhances compositional fidelity and image quality across various styles and concept sets, outperforming state-of-the-art baselines in multi-concept generation tasks. Code will be released.
title MultLFG: Training-free Multi-LoRA composition using Frequency-domain Guidance
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.20525