DreamOmni: Unified Image Generation and Editing

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Xia, Bin, Zhang, Yuechen, Li, Jingyao, Wang, Chengyao, Wang, Yitong, Wu, Xinglong, Yu, Bei, Jia, Jiaya
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909819449901056
author Xia, Bin
Zhang, Yuechen
Li, Jingyao
Wang, Chengyao
Wang, Yitong
Wu, Xinglong
Yu, Bei
Jia, Jiaya
author_facet Xia, Bin
Zhang, Yuechen
Li, Jingyao
Wang, Chengyao
Wang, Yitong
Wu, Xinglong
Yu, Bei
Jia, Jiaya
contents Currently, the success of large language models (LLMs) illustrates that a unified multitasking approach can significantly enhance model usability, streamline deployment, and foster synergistic benefits across different tasks. However, in computer vision, while text-to-image (T2I) models have significantly improved generation quality through scaling up, their framework design did not initially consider how to unify with downstream tasks, such as various types of editing. To address this, we introduce DreamOmni, a unified model for image generation and editing. We begin by analyzing existing frameworks and the requirements of downstream tasks, proposing a unified framework that integrates both T2I models and various editing tasks. Furthermore, another key challenge is the efficient creation of high-quality editing data, particularly for instruction-based and drag-based editing. To this end, we develop a synthetic data pipeline using sticker-like elements to synthesize accurate, high-quality datasets efficiently, which enables editing data scaling up for unified model training. For training, DreamOmni jointly trains T2I generation and downstream tasks. T2I training enhances the model's understanding of specific concepts and improves generation quality, while editing training helps the model grasp the nuances of the editing task. This collaboration significantly boosts editing performance. Extensive experiments confirm the effectiveness of DreamOmni. The code and model will be released.
format Preprint
id arxiv_https___arxiv_org_abs_2412_17098
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle DreamOmni: Unified Image Generation and Editing
Xia, Bin
Zhang, Yuechen
Li, Jingyao
Wang, Chengyao
Wang, Yitong
Wu, Xinglong
Yu, Bei
Jia, Jiaya
Computer Vision and Pattern Recognition
Currently, the success of large language models (LLMs) illustrates that a unified multitasking approach can significantly enhance model usability, streamline deployment, and foster synergistic benefits across different tasks. However, in computer vision, while text-to-image (T2I) models have significantly improved generation quality through scaling up, their framework design did not initially consider how to unify with downstream tasks, such as various types of editing. To address this, we introduce DreamOmni, a unified model for image generation and editing. We begin by analyzing existing frameworks and the requirements of downstream tasks, proposing a unified framework that integrates both T2I models and various editing tasks. Furthermore, another key challenge is the efficient creation of high-quality editing data, particularly for instruction-based and drag-based editing. To this end, we develop a synthetic data pipeline using sticker-like elements to synthesize accurate, high-quality datasets efficiently, which enables editing data scaling up for unified model training. For training, DreamOmni jointly trains T2I generation and downstream tasks. T2I training enhances the model's understanding of specific concepts and improves generation quality, while editing training helps the model grasp the nuances of the editing task. This collaboration significantly boosts editing performance. Extensive experiments confirm the effectiveness of DreamOmni. The code and model will be released.
title DreamOmni: Unified Image Generation and Editing
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.17098