G4Seg: Generation for Inexact Segmentation Refinement with Diffusion Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Tianjiao, Zhang, Fei, Yao, Jiangchao, Zhang, Ya, Wang, Yanfeng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912409455689728
author Zhang, Tianjiao
Zhang, Fei
Yao, Jiangchao
Zhang, Ya
Wang, Yanfeng
author_facet Zhang, Tianjiao
Zhang, Fei
Yao, Jiangchao
Zhang, Ya
Wang, Yanfeng
contents This paper considers the problem of utilizing a large-scale text-to-image diffusion model to tackle the challenging Inexact Segmentation (IS) task. Unlike traditional approaches that rely heavily on discriminative-model-based paradigms or dense visual representations derived from internal attention mechanisms, our method focuses on the intrinsic generative priors in Stable Diffusion~(SD). Specifically, we exploit the pattern discrepancies between original images and mask-conditional generated images to facilitate a coarse-to-fine segmentation refinement by establishing a semantic correspondence alignment and updating the foreground probability. Comprehensive quantitative and qualitative experiments validate the effectiveness and superiority of our plug-and-play design, underscoring the potential of leveraging generation discrepancies to model dense representations and encouraging further exploration of generative approaches for solving discriminative tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2506_01539
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle G4Seg: Generation for Inexact Segmentation Refinement with Diffusion Models
Zhang, Tianjiao
Zhang, Fei
Yao, Jiangchao
Zhang, Ya
Wang, Yanfeng
Computer Vision and Pattern Recognition
Artificial Intelligence
This paper considers the problem of utilizing a large-scale text-to-image diffusion model to tackle the challenging Inexact Segmentation (IS) task. Unlike traditional approaches that rely heavily on discriminative-model-based paradigms or dense visual representations derived from internal attention mechanisms, our method focuses on the intrinsic generative priors in Stable Diffusion~(SD). Specifically, we exploit the pattern discrepancies between original images and mask-conditional generated images to facilitate a coarse-to-fine segmentation refinement by establishing a semantic correspondence alignment and updating the foreground probability. Comprehensive quantitative and qualitative experiments validate the effectiveness and superiority of our plug-and-play design, underscoring the potential of leveraging generation discrepancies to model dense representations and encouraging further exploration of generative approaches for solving discriminative tasks.
title G4Seg: Generation for Inexact Segmentation Refinement with Diffusion Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2506.01539