UniCustom: Unified Visual Conditioning for Multi-Reference Image Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Yiyan, Wang, Qiulin, Wang, Wenjie, Mao, Yunyao, Wang, Xintao, Wan, Pengfei, Gai, Kun, Feng, Fuli
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910215365984256
author Xu, Yiyan
Wang, Qiulin
Wang, Wenjie
Mao, Yunyao
Wang, Xintao
Wan, Pengfei
Gai, Kun
Feng, Fuli
author_facet Xu, Yiyan
Wang, Qiulin
Wang, Wenjie
Mao, Yunyao
Wang, Xintao
Wan, Pengfei
Gai, Kun
Feng, Fuli
contents Multi-reference image generation aims to synthesize images from textual instructions while faithfully preserving subject identities from multiple reference images. Existing VLM-enhanced diffusion models commonly rely on decoupled visual conditioning: semantic ViT features are processed by the VLM for instruction understanding, whereas appearance-rich VAE features are injected later into the diffusion backbone. Despite its intuitive design, this separation makes it difficult for the model to associate each semantically grounded subject with visual details from the correct reference image. As a result, the model may recognize which subject is being referred to, but fail to preserve its identity and fine-grained appearance, leading to attribute leakage and cross-reference confusion in complex multi-reference settings. To address this issue, we propose UniCustom, a unified visual conditioning framework that fuses ViT and VAE features before VLM encoding. This early fusion exposes the VLM to both semantic cues and appearance-rich details, enabling its hidden states to jointly encode the referred subject and corresponding visual appearance with only a lightweight linear fusion layer. To learn such unified representations, we adopt a two-stage training strategy: reconstruction-oriented pretraining that preserves reference-specific appearance details in the fused hidden states, followed by supervised finetuning on single- and multi-reference generation tasks. We further introduce a slot-wise binding regularization that encourages each image slot to preserve low-level details of its corresponding reference, thereby reducing cross-reference entanglement. Experiments on two multi-reference generation benchmarks demonstrate that UniCustom consistently improves subject consistency, instruction following, and compositional fidelity over strong baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2605_12088
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle UniCustom: Unified Visual Conditioning for Multi-Reference Image Generation
Xu, Yiyan
Wang, Qiulin
Wang, Wenjie
Mao, Yunyao
Wang, Xintao
Wan, Pengfei
Gai, Kun
Feng, Fuli
Computer Vision and Pattern Recognition
Multi-reference image generation aims to synthesize images from textual instructions while faithfully preserving subject identities from multiple reference images. Existing VLM-enhanced diffusion models commonly rely on decoupled visual conditioning: semantic ViT features are processed by the VLM for instruction understanding, whereas appearance-rich VAE features are injected later into the diffusion backbone. Despite its intuitive design, this separation makes it difficult for the model to associate each semantically grounded subject with visual details from the correct reference image. As a result, the model may recognize which subject is being referred to, but fail to preserve its identity and fine-grained appearance, leading to attribute leakage and cross-reference confusion in complex multi-reference settings. To address this issue, we propose UniCustom, a unified visual conditioning framework that fuses ViT and VAE features before VLM encoding. This early fusion exposes the VLM to both semantic cues and appearance-rich details, enabling its hidden states to jointly encode the referred subject and corresponding visual appearance with only a lightweight linear fusion layer. To learn such unified representations, we adopt a two-stage training strategy: reconstruction-oriented pretraining that preserves reference-specific appearance details in the fused hidden states, followed by supervised finetuning on single- and multi-reference generation tasks. We further introduce a slot-wise binding regularization that encourages each image slot to preserve low-level details of its corresponding reference, thereby reducing cross-reference entanglement. Experiments on two multi-reference generation benchmarks demonstrate that UniCustom consistently improves subject consistency, instruction following, and compositional fidelity over strong baselines.
title UniCustom: Unified Visual Conditioning for Multi-Reference Image Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.12088