ScaleWeaver: Weaving Efficient Controllable T2I Generation with Multi-Scale Reference Attention

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Liu, Keli, Wang, Zhendong, Zhou, Wengang, Xu, Shaodong, Dong, Ruixiao, Li, Houqiang
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909850549616640
author Liu, Keli
Wang, Zhendong
Zhou, Wengang
Xu, Shaodong
Dong, Ruixiao
Li, Houqiang
author_facet Liu, Keli
Wang, Zhendong
Zhou, Wengang
Xu, Shaodong
Dong, Ruixiao
Li, Houqiang
contents Text-to-image generation with visual autoregressive~(VAR) models has recently achieved impressive advances in generation fidelity and inference efficiency. While control mechanisms have been explored for diffusion models, enabling precise and flexible control within VAR paradigm remains underexplored. To bridge this critical gap, in this paper, we introduce ScaleWeaver, a novel framework designed to achieve high-fidelity, controllable generation upon advanced VAR models through parameter-efficient fine-tuning. The core module in ScaleWeaver is the improved MMDiT block with the proposed Reference Attention module, which efficiently and effectively incorporates conditional information. Different from MM Attention, the proposed Reference Attention module discards the unnecessary attention from image$\rightarrow$condition, reducing computational cost while stabilizing control injection. Besides, it strategically emphasizes parameter reuse, leveraging the capability of the VAR backbone itself with a few introduced parameters to process control information, and equipping a zero-initialized linear projection to ensure that control signals are incorporated effectively without disrupting the generative capability of the base model. Extensive experiments show that ScaleWeaver delivers high-quality generation and precise control while attaining superior efficiency over diffusion-based methods, making ScaleWeaver a practical and effective solution for controllable text-to-image generation within the visual autoregressive paradigm. Code and models will be released.
format Preprint
id arxiv_https___arxiv_org_abs_2510_14882
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ScaleWeaver: Weaving Efficient Controllable T2I Generation with Multi-Scale Reference Attention
Liu, Keli
Wang, Zhendong
Zhou, Wengang
Xu, Shaodong
Dong, Ruixiao
Li, Houqiang
Computer Vision and Pattern Recognition
Text-to-image generation with visual autoregressive~(VAR) models has recently achieved impressive advances in generation fidelity and inference efficiency. While control mechanisms have been explored for diffusion models, enabling precise and flexible control within VAR paradigm remains underexplored. To bridge this critical gap, in this paper, we introduce ScaleWeaver, a novel framework designed to achieve high-fidelity, controllable generation upon advanced VAR models through parameter-efficient fine-tuning. The core module in ScaleWeaver is the improved MMDiT block with the proposed Reference Attention module, which efficiently and effectively incorporates conditional information. Different from MM Attention, the proposed Reference Attention module discards the unnecessary attention from image$\rightarrow$condition, reducing computational cost while stabilizing control injection. Besides, it strategically emphasizes parameter reuse, leveraging the capability of the VAR backbone itself with a few introduced parameters to process control information, and equipping a zero-initialized linear projection to ensure that control signals are incorporated effectively without disrupting the generative capability of the base model. Extensive experiments show that ScaleWeaver delivers high-quality generation and precise control while attaining superior efficiency over diffusion-based methods, making ScaleWeaver a practical and effective solution for controllable text-to-image generation within the visual autoregressive paradigm. Code and models will be released.
title ScaleWeaver: Weaving Efficient Controllable T2I Generation with Multi-Scale Reference Attention
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.14882