DreamOmni3: Scribble-based Editing and Generation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Xia, Bin, Peng, Bohao, Liu, Jiyang, Wu, Sitong, Li, Jingyao, Huang, Junjia, Zhao, Xu, Wang, Yitong, Chu, Ruihang, Yu, Bei, Jia, Jiaya
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911340955697152
author Xia, Bin
Peng, Bohao
Liu, Jiyang
Wu, Sitong
Li, Jingyao
Huang, Junjia
Zhao, Xu
Wang, Yitong
Chu, Ruihang
Yu, Bei
Jia, Jiaya
author_facet Xia, Bin
Peng, Bohao
Liu, Jiyang
Wu, Sitong
Li, Jingyao
Huang, Junjia
Zhao, Xu
Wang, Yitong
Chu, Ruihang
Yu, Bei
Jia, Jiaya
contents Recently unified generation and editing models have achieved remarkable success with their impressive performance. These models rely mainly on text prompts for instruction-based editing and generation, but language often fails to capture users intended edit locations and fine-grained visual details. To this end, we propose two tasks: scribble-based editing and generation, that enables more flexible creation on graphical user interface (GUI) combining user textual, images, and freehand sketches. We introduce DreamOmni3, tackling two challenges: data creation and framework design. Our data synthesis pipeline includes two parts: scribble-based editing and generation. For scribble-based editing, we define four tasks: scribble and instruction-based editing, scribble and multimodal instruction-based editing, image fusion, and doodle editing. Based on DreamOmni2 dataset, we extract editable regions and overlay hand-drawn boxes, circles, doodles or cropped image to construct training data. For scribble-based generation, we define three tasks: scribble and instruction-based generation, scribble and multimodal instruction-based generation, and doodle generation, following similar data creation pipelines. For the framework, instead of using binary masks, which struggle with complex edits involving multiple scribbles, images, and instructions, we propose a joint input scheme that feeds both the original and scribbled source images into the model, using different colors to distinguish regions and simplify processing. By applying the same index and position encodings to both images, the model can precisely localize scribbled regions while maintaining accurate editing. Finally, we establish comprehensive benchmarks for these tasks to promote further research. Experimental results demonstrate that DreamOmni3 achieves outstanding performance, and models and code will be publicly released.
format Preprint
id arxiv_https___arxiv_org_abs_2512_22525
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DreamOmni3: Scribble-based Editing and Generation
Xia, Bin
Peng, Bohao
Liu, Jiyang
Wu, Sitong
Li, Jingyao
Huang, Junjia
Zhao, Xu
Wang, Yitong
Chu, Ruihang
Yu, Bei
Jia, Jiaya
Computer Vision and Pattern Recognition
Recently unified generation and editing models have achieved remarkable success with their impressive performance. These models rely mainly on text prompts for instruction-based editing and generation, but language often fails to capture users intended edit locations and fine-grained visual details. To this end, we propose two tasks: scribble-based editing and generation, that enables more flexible creation on graphical user interface (GUI) combining user textual, images, and freehand sketches. We introduce DreamOmni3, tackling two challenges: data creation and framework design. Our data synthesis pipeline includes two parts: scribble-based editing and generation. For scribble-based editing, we define four tasks: scribble and instruction-based editing, scribble and multimodal instruction-based editing, image fusion, and doodle editing. Based on DreamOmni2 dataset, we extract editable regions and overlay hand-drawn boxes, circles, doodles or cropped image to construct training data. For scribble-based generation, we define three tasks: scribble and instruction-based generation, scribble and multimodal instruction-based generation, and doodle generation, following similar data creation pipelines. For the framework, instead of using binary masks, which struggle with complex edits involving multiple scribbles, images, and instructions, we propose a joint input scheme that feeds both the original and scribbled source images into the model, using different colors to distinguish regions and simplify processing. By applying the same index and position encodings to both images, the model can precisely localize scribbled regions while maintaining accurate editing. Finally, we establish comprehensive benchmarks for these tasks to promote further research. Experimental results demonstrate that DreamOmni3 achieves outstanding performance, and models and code will be publicly released.
title DreamOmni3: Scribble-based Editing and Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.22525