DreamOmni2: Multimodal Instruction-based Editing and Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xia, Bin, Peng, Bohao, Zhang, Yuechen, Huang, Junjia, Liu, Jiyang, Li, Jingyao, Tan, Haoru, Wu, Sitong, Wang, Chengyao, Wang, Yitong, Wu, Xinglong, Yu, Bei, Jia, Jiaya
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909831015694336
author Xia, Bin
Peng, Bohao
Zhang, Yuechen
Huang, Junjia
Liu, Jiyang
Li, Jingyao
Tan, Haoru
Wu, Sitong
Wang, Chengyao
Wang, Yitong
Wu, Xinglong
Yu, Bei
Jia, Jiaya
author_facet Xia, Bin
Peng, Bohao
Zhang, Yuechen
Huang, Junjia
Liu, Jiyang
Li, Jingyao
Tan, Haoru
Wu, Sitong
Wang, Chengyao
Wang, Yitong
Wu, Xinglong
Yu, Bei
Jia, Jiaya
contents Recent advancements in instruction-based image editing and subject-driven generation have garnered significant attention, yet both tasks still face limitations in meeting practical user needs. Instruction-based editing relies solely on language instructions, which often fail to capture specific editing details, making reference images necessary. Meanwhile, subject-driven generation is limited to combining concrete objects or people, overlooking broader, abstract concepts. To address these challenges, we propose two novel tasks: multimodal instruction-based editing and generation. These tasks support both text and image instructions and extend the scope to include both concrete and abstract concepts, greatly enhancing their practical applications. We introduce DreamOmni2, tackling two primary challenges: data creation and model framework design. Our data synthesis pipeline consists of three steps: (1) using a feature mixing method to create extraction data for both abstract and concrete concepts, (2) generating multimodal instruction-based editing training data using the editing and extraction models, and (3) further applying the extraction model to create training data for multimodal instruction-based editing. For the framework, to handle multi-image input, we propose an index encoding and position encoding shift scheme, which helps the model distinguish images and avoid pixel confusion. Additionally, we introduce joint training with the VLM and our generation/editing model to better process complex instructions. In addition, we have proposed comprehensive benchmarks for these two new tasks to drive their development. Experiments show that DreamOmni2 has achieved impressive results. Models and codes will be released.
format Preprint
id arxiv_https___arxiv_org_abs_2510_06679
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DreamOmni2: Multimodal Instruction-based Editing and Generation
Xia, Bin
Peng, Bohao
Zhang, Yuechen
Huang, Junjia
Liu, Jiyang
Li, Jingyao
Tan, Haoru
Wu, Sitong
Wang, Chengyao
Wang, Yitong
Wu, Xinglong
Yu, Bei
Jia, Jiaya
Computer Vision and Pattern Recognition
Recent advancements in instruction-based image editing and subject-driven generation have garnered significant attention, yet both tasks still face limitations in meeting practical user needs. Instruction-based editing relies solely on language instructions, which often fail to capture specific editing details, making reference images necessary. Meanwhile, subject-driven generation is limited to combining concrete objects or people, overlooking broader, abstract concepts. To address these challenges, we propose two novel tasks: multimodal instruction-based editing and generation. These tasks support both text and image instructions and extend the scope to include both concrete and abstract concepts, greatly enhancing their practical applications. We introduce DreamOmni2, tackling two primary challenges: data creation and model framework design. Our data synthesis pipeline consists of three steps: (1) using a feature mixing method to create extraction data for both abstract and concrete concepts, (2) generating multimodal instruction-based editing training data using the editing and extraction models, and (3) further applying the extraction model to create training data for multimodal instruction-based editing. For the framework, to handle multi-image input, we propose an index encoding and position encoding shift scheme, which helps the model distinguish images and avoid pixel confusion. Additionally, we introduce joint training with the VLM and our generation/editing model to better process complex instructions. In addition, we have proposed comprehensive benchmarks for these two new tasks to drive their development. Experiments show that DreamOmni2 has achieved impressive results. Models and codes will be released.
title DreamOmni2: Multimodal Instruction-based Editing and Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.06679