FlowDC: Flow-Based Decoupling-Decay for Complex Image Editing

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Jiang, Yilei, Wang, Zhen, Wang, Yanghao, Yu, Jun, Zhuang, Yueting, Xiao, Jun, Chen, Long
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914197386821632
author Jiang, Yilei
Wang, Zhen
Wang, Yanghao
Yu, Jun
Zhuang, Yueting
Xiao, Jun
Chen, Long
author_facet Jiang, Yilei
Wang, Zhen
Wang, Yanghao
Yu, Jun
Zhuang, Yueting
Xiao, Jun
Chen, Long
contents With the surge of pre-trained text-to-image flow matching models, text-based image editing performance has gained remarkable improvement, especially for \underline{simple editing} that only contains a single editing target. To satisfy the exploding editing requirements, the \underline{complex editing} which contains multiple editing targets has posed as a more challenging task. However, current complex editing solutions: single-round and multi-round editing are limited by long text following and cumulative inconsistency, respectively. Thus, they struggle to strike a balance between semantic alignment and source consistency. In this paper, we propose \textbf{FlowDC}, which decouples the complex editing into multiple sub-editing effects and superposes them in parallel during the editing process. Meanwhile, we observed that the velocity quantity that is orthogonal to the editing displacement harms the source structure preserving. Thus, we decompose the velocity and decay the orthogonal part for better source consistency. To evaluate the effectiveness of complex editing settings, we construct a complex editing benchmark: Complex-PIE-Bench. On two benchmarks, FlowDC shows superior results compared with existing methods. We also detail the ablations of our module designs.
format Preprint
id arxiv_https___arxiv_org_abs_2512_11395
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FlowDC: Flow-Based Decoupling-Decay for Complex Image Editing
Jiang, Yilei
Wang, Zhen
Wang, Yanghao
Yu, Jun
Zhuang, Yueting
Xiao, Jun
Chen, Long
Computer Vision and Pattern Recognition
With the surge of pre-trained text-to-image flow matching models, text-based image editing performance has gained remarkable improvement, especially for \underline{simple editing} that only contains a single editing target. To satisfy the exploding editing requirements, the \underline{complex editing} which contains multiple editing targets has posed as a more challenging task. However, current complex editing solutions: single-round and multi-round editing are limited by long text following and cumulative inconsistency, respectively. Thus, they struggle to strike a balance between semantic alignment and source consistency. In this paper, we propose \textbf{FlowDC}, which decouples the complex editing into multiple sub-editing effects and superposes them in parallel during the editing process. Meanwhile, we observed that the velocity quantity that is orthogonal to the editing displacement harms the source structure preserving. Thus, we decompose the velocity and decay the orthogonal part for better source consistency. To evaluate the effectiveness of complex editing settings, we construct a complex editing benchmark: Complex-PIE-Bench. On two benchmarks, FlowDC shows superior results compared with existing methods. We also detail the ablations of our module designs.
title FlowDC: Flow-Based Decoupling-Decay for Complex Image Editing
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.11395