Describe, Don't Dictate: Semantic Image Editing with Natural Language Intent

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Ci, En, Guan, Shanyan, Ge, Yanhao, Zhang, Yilin, Li, Wei, Zhang, Zhenyu, Yang, Jian, Tai, Ying
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918131943866368
author Ci, En
Guan, Shanyan
Ge, Yanhao
Zhang, Yilin
Li, Wei
Zhang, Zhenyu
Yang, Jian
Tai, Ying
author_facet Ci, En
Guan, Shanyan
Ge, Yanhao
Zhang, Yilin
Li, Wei
Zhang, Zhenyu
Yang, Jian
Tai, Ying
contents Despite the progress in text-to-image generation, semantic image editing remains a challenge. Inversion-based algorithms unavoidably introduce reconstruction errors, while instruction-based models mainly suffer from limited dataset quality and scale. To address these problems, we propose a descriptive-prompt-based editing framework, named DescriptiveEdit. The core idea is to re-frame `instruction-based image editing' as `reference-image-based text-to-image generation', which preserves the generative power of well-trained Text-to-Image models without architectural modifications or inversion. Specifically, taking the reference image and a prompt as input, we introduce a Cross-Attentive UNet, which newly adds attention bridges to inject reference image features into the prompt-to-edit-image generation process. Owing to its text-to-image nature, DescriptiveEdit overcomes limitations in instruction dataset quality, integrates seamlessly with ControlNet, IP-Adapter, and other extensions, and is more scalable. Experiments on the Emu Edit benchmark show it improves editing accuracy and consistency.
format Preprint
id arxiv_https___arxiv_org_abs_2508_20505
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Describe, Don't Dictate: Semantic Image Editing with Natural Language Intent
Ci, En
Guan, Shanyan
Ge, Yanhao
Zhang, Yilin
Li, Wei
Zhang, Zhenyu
Yang, Jian
Tai, Ying
Computer Vision and Pattern Recognition
Despite the progress in text-to-image generation, semantic image editing remains a challenge. Inversion-based algorithms unavoidably introduce reconstruction errors, while instruction-based models mainly suffer from limited dataset quality and scale. To address these problems, we propose a descriptive-prompt-based editing framework, named DescriptiveEdit. The core idea is to re-frame `instruction-based image editing' as `reference-image-based text-to-image generation', which preserves the generative power of well-trained Text-to-Image models without architectural modifications or inversion. Specifically, taking the reference image and a prompt as input, we introduce a Cross-Attentive UNet, which newly adds attention bridges to inject reference image features into the prompt-to-edit-image generation process. Owing to its text-to-image nature, DescriptiveEdit overcomes limitations in instruction dataset quality, integrates seamlessly with ControlNet, IP-Adapter, and other extensions, and is more scalable. Experiments on the Emu Edit benchmark show it improves editing accuracy and consistency.
title Describe, Don't Dictate: Semantic Image Editing with Natural Language Intent
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.20505