AnyMS: Bottom-up Attention Decoupling for Layout-guided and Training-free Multi-subject Customization

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Yu, Binhe, Wang, Zhen, Li, Kexin, Yuan, Yuqian, Zhang, Wenqiao, Chen, Long, Li, Juncheng, Xiao, Jun, Zhuang, Yueting
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909979655536640
author Yu, Binhe
Wang, Zhen
Li, Kexin
Yuan, Yuqian
Zhang, Wenqiao
Chen, Long
Li, Juncheng
Xiao, Jun
Zhuang, Yueting
author_facet Yu, Binhe
Wang, Zhen
Li, Kexin
Yuan, Yuqian
Zhang, Wenqiao
Chen, Long
Li, Juncheng
Xiao, Jun
Zhuang, Yueting
contents Multi-subject customization aims to synthesize multiple user-specified subjects into a coherent image. To address issues such as subjects missing or conflicts, recent works incorporate layout guidance to provide explicit spatial constraints. However, existing methods still struggle to balance three critical objectives: text alignment, subject identity preservation, and layout control, while the reliance on additional training further limits their scalability and efficiency. In this paper, we present AnyMS, a novel training-free framework for layout-guided multi-subject customization. AnyMS leverages three input conditions: text prompt, subject images, and layout constraints, and introduces a bottom-up dual-level attention decoupling mechanism to harmonize their integration during generation. Specifically, global decoupling separates cross-attention between textual and visual conditions to ensure text alignment. Local decoupling confines each subject's attention to its designated area, which prevents subject conflicts and thus guarantees identity preservation and layout control. Moreover, AnyMS employs pre-trained image adapters to extract subject-specific features aligned with the diffusion model, removing the need for subject learning or adapter tuning. Extensive experiments demonstrate that AnyMS achieves state-of-the-art performance, supporting complex compositions and scaling to a larger number of subjects.
format Preprint
id arxiv_https___arxiv_org_abs_2512_23537
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AnyMS: Bottom-up Attention Decoupling for Layout-guided and Training-free Multi-subject Customization
Yu, Binhe
Wang, Zhen
Li, Kexin
Yuan, Yuqian
Zhang, Wenqiao
Chen, Long
Li, Juncheng
Xiao, Jun
Zhuang, Yueting
Computer Vision and Pattern Recognition
Artificial Intelligence
Multi-subject customization aims to synthesize multiple user-specified subjects into a coherent image. To address issues such as subjects missing or conflicts, recent works incorporate layout guidance to provide explicit spatial constraints. However, existing methods still struggle to balance three critical objectives: text alignment, subject identity preservation, and layout control, while the reliance on additional training further limits their scalability and efficiency. In this paper, we present AnyMS, a novel training-free framework for layout-guided multi-subject customization. AnyMS leverages three input conditions: text prompt, subject images, and layout constraints, and introduces a bottom-up dual-level attention decoupling mechanism to harmonize their integration during generation. Specifically, global decoupling separates cross-attention between textual and visual conditions to ensure text alignment. Local decoupling confines each subject's attention to its designated area, which prevents subject conflicts and thus guarantees identity preservation and layout control. Moreover, AnyMS employs pre-trained image adapters to extract subject-specific features aligned with the diffusion model, removing the need for subject learning or adapter tuning. Extensive experiments demonstrate that AnyMS achieves state-of-the-art performance, supporting complex compositions and scaling to a larger number of subjects.
title AnyMS: Bottom-up Attention Decoupling for Layout-guided and Training-free Multi-subject Customization
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2512.23537