Edit2Perceive: Image Editing Diffusion Models Are Strong Dense Perceivers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shi, Yiqing, Song, Yiren, Shou, Mike Zheng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915634784239616
author Shi, Yiqing
Song, Yiren
Shou, Mike Zheng
author_facet Shi, Yiqing
Song, Yiren
Shou, Mike Zheng
contents Recent advances in diffusion transformers have shown remarkable generalization in visual synthesis, yet most dense perception methods still rely on text-to-image (T2I) generators designed for stochastic generation. We revisit this paradigm and show that image editing diffusion models are inherently image-to-image consistent, providing a more suitable foundation for dense perception task. We introduce Edit2Perceive, a unified diffusion framework that adapts editing models for depth, normal, and matting. Built upon the FLUX.1 Kontext architecture, our approach employs full-parameter fine-tuning and a pixel-space consistency loss to enforce structure-preserving refinement across intermediate denoising states. Moreover, our single-step deterministic inference yields up to faster runtime while training on relatively small datasets. Extensive experiments demonstrate comprehensive state-of-the-art results across all three tasks, revealing the strong potential of editing-oriented diffusion transformers for geometry-aware perception.
format Preprint
id arxiv_https___arxiv_org_abs_2511_18673
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Edit2Perceive: Image Editing Diffusion Models Are Strong Dense Perceivers
Shi, Yiqing
Song, Yiren
Shou, Mike Zheng
Computer Vision and Pattern Recognition
Recent advances in diffusion transformers have shown remarkable generalization in visual synthesis, yet most dense perception methods still rely on text-to-image (T2I) generators designed for stochastic generation. We revisit this paradigm and show that image editing diffusion models are inherently image-to-image consistent, providing a more suitable foundation for dense perception task. We introduce Edit2Perceive, a unified diffusion framework that adapts editing models for depth, normal, and matting. Built upon the FLUX.1 Kontext architecture, our approach employs full-parameter fine-tuning and a pixel-space consistency loss to enforce structure-preserving refinement across intermediate denoising states. Moreover, our single-step deterministic inference yields up to faster runtime while training on relatively small datasets. Extensive experiments demonstrate comprehensive state-of-the-art results across all three tasks, revealing the strong potential of editing-oriented diffusion transformers for geometry-aware perception.
title Edit2Perceive: Image Editing Diffusion Models Are Strong Dense Perceivers
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.18673