MultiRef: Controllable Image Generation with Multiple Visual References

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Ruoxi, Chen, Dongping, Wu, Siyuan, Wang, Sinan, Lang, Shiyun, Sushko, Petr, Jiang, Gaoyang, Wan, Yao, Krishna, Ranjay
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909753207160832
author Chen, Ruoxi
Chen, Dongping
Wu, Siyuan
Wang, Sinan
Lang, Shiyun
Sushko, Petr
Jiang, Gaoyang
Wan, Yao
Krishna, Ranjay
author_facet Chen, Ruoxi
Chen, Dongping
Wu, Siyuan
Wang, Sinan
Lang, Shiyun
Sushko, Petr
Jiang, Gaoyang
Wan, Yao
Krishna, Ranjay
contents Visual designers naturally draw inspiration from multiple visual references, combining diverse elements and aesthetic principles to create artwork. However, current image generative frameworks predominantly rely on single-source inputs -- either text prompts or individual reference images. In this paper, we focus on the task of controllable image generation using multiple visual references. We introduce MultiRef-bench, a rigorous evaluation framework comprising 990 synthetic and 1,000 real-world samples that require incorporating visual content from multiple reference images. The synthetic samples are synthetically generated through our data engine RefBlend, with 10 reference types and 33 reference combinations. Based on RefBlend, we further construct a dataset MultiRef containing 38k high-quality images to facilitate further research. Our experiments across three interleaved image-text models (i.e., OmniGen, ACE, and Show-o) and six agentic frameworks (e.g., ChatDiT and LLM + SD) reveal that even state-of-the-art systems struggle with multi-reference conditioning, with the best model OmniGen achieving only 66.6% in synthetic samples and 79.0% in real-world cases on average compared to the golden answer. These findings provide valuable directions for developing more flexible and human-like creative tools that can effectively integrate multiple sources of visual inspiration. The dataset is publicly available at: https://multiref.github.io/.
format Preprint
id arxiv_https___arxiv_org_abs_2508_06905
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MultiRef: Controllable Image Generation with Multiple Visual References
Chen, Ruoxi
Chen, Dongping
Wu, Siyuan
Wang, Sinan
Lang, Shiyun
Sushko, Petr
Jiang, Gaoyang
Wan, Yao
Krishna, Ranjay
Computer Vision and Pattern Recognition
Visual designers naturally draw inspiration from multiple visual references, combining diverse elements and aesthetic principles to create artwork. However, current image generative frameworks predominantly rely on single-source inputs -- either text prompts or individual reference images. In this paper, we focus on the task of controllable image generation using multiple visual references. We introduce MultiRef-bench, a rigorous evaluation framework comprising 990 synthetic and 1,000 real-world samples that require incorporating visual content from multiple reference images. The synthetic samples are synthetically generated through our data engine RefBlend, with 10 reference types and 33 reference combinations. Based on RefBlend, we further construct a dataset MultiRef containing 38k high-quality images to facilitate further research. Our experiments across three interleaved image-text models (i.e., OmniGen, ACE, and Show-o) and six agentic frameworks (e.g., ChatDiT and LLM + SD) reveal that even state-of-the-art systems struggle with multi-reference conditioning, with the best model OmniGen achieving only 66.6% in synthetic samples and 79.0% in real-world cases on average compared to the golden answer. These findings provide valuable directions for developing more flexible and human-like creative tools that can effectively integrate multiple sources of visual inspiration. The dataset is publicly available at: https://multiref.github.io/.
title MultiRef: Controllable Image Generation with Multiple Visual References
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.06905