UNIC-Adapter: Unified Image-instruction Adapter with Multi-modal Transformer for Image Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Duan, Lunhao, Zhao, Shanshan, Yan, Wenjun, Li, Yinglun, Chen, Qing-Guo, Xu, Zhao, Luo, Weihua, Zhang, Kaifu, Gong, Mingming, Xia, Gui-Song
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909441140457472
author Duan, Lunhao
Zhao, Shanshan
Yan, Wenjun
Li, Yinglun
Chen, Qing-Guo
Xu, Zhao
Luo, Weihua
Zhang, Kaifu
Gong, Mingming
Xia, Gui-Song
author_facet Duan, Lunhao
Zhao, Shanshan
Yan, Wenjun
Li, Yinglun
Chen, Qing-Guo
Xu, Zhao
Luo, Weihua
Zhang, Kaifu
Gong, Mingming
Xia, Gui-Song
contents Recently, text-to-image generation models have achieved remarkable advancements, particularly with diffusion models facilitating high-quality image synthesis from textual descriptions. However, these models often struggle with achieving precise control over pixel-level layouts, object appearances, and global styles when using text prompts alone. To mitigate this issue, previous works introduce conditional images as auxiliary inputs for image generation, enhancing control but typically necessitating specialized models tailored to different types of reference inputs. In this paper, we explore a new approach to unify controllable generation within a single framework. Specifically, we propose the unified image-instruction adapter (UNIC-Adapter) built on the Multi-Modal-Diffusion Transformer architecture, to enable flexible and controllable generation across diverse conditions without the need for multiple specialized models. Our UNIC-Adapter effectively extracts multi-modal instruction information by incorporating both conditional images and task instructions, injecting this information into the image generation process through a cross-attention mechanism enhanced by Rotary Position Embedding. Experimental results across a variety of tasks, including pixel-level spatial control, subject-driven image generation, and style-image-based image synthesis, demonstrate the effectiveness of our UNIC-Adapter in unified controllable image generation.
format Preprint
id arxiv_https___arxiv_org_abs_2412_18928
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle UNIC-Adapter: Unified Image-instruction Adapter with Multi-modal Transformer for Image Generation
Duan, Lunhao
Zhao, Shanshan
Yan, Wenjun
Li, Yinglun
Chen, Qing-Guo
Xu, Zhao
Luo, Weihua
Zhang, Kaifu
Gong, Mingming
Xia, Gui-Song
Computer Vision and Pattern Recognition
Machine Learning
Recently, text-to-image generation models have achieved remarkable advancements, particularly with diffusion models facilitating high-quality image synthesis from textual descriptions. However, these models often struggle with achieving precise control over pixel-level layouts, object appearances, and global styles when using text prompts alone. To mitigate this issue, previous works introduce conditional images as auxiliary inputs for image generation, enhancing control but typically necessitating specialized models tailored to different types of reference inputs. In this paper, we explore a new approach to unify controllable generation within a single framework. Specifically, we propose the unified image-instruction adapter (UNIC-Adapter) built on the Multi-Modal-Diffusion Transformer architecture, to enable flexible and controllable generation across diverse conditions without the need for multiple specialized models. Our UNIC-Adapter effectively extracts multi-modal instruction information by incorporating both conditional images and task instructions, injecting this information into the image generation process through a cross-attention mechanism enhanced by Rotary Position Embedding. Experimental results across a variety of tasks, including pixel-level spatial control, subject-driven image generation, and style-image-based image synthesis, demonstrate the effectiveness of our UNIC-Adapter in unified controllable image generation.
title UNIC-Adapter: Unified Image-instruction Adapter with Multi-modal Transformer for Image Generation
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2412.18928