SGDiff: Scene Graph Guided Diffusion Model for Image Collaborative SegCaptioning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Xu, Yuan, Jin, Zhang, Hanwang, Zhong, Guojin, Zang, Yongsheng, Lin, Jiacheng, Li, Zhiyong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917116525936640
author Zhang, Xu
Yuan, Jin
Zhang, Hanwang
Zhong, Guojin
Zang, Yongsheng
Lin, Jiacheng
Li, Zhiyong
author_facet Zhang, Xu
Yuan, Jin
Zhang, Hanwang
Zhong, Guojin
Zang, Yongsheng
Lin, Jiacheng
Li, Zhiyong
contents Controllable image semantic understanding tasks, such as captioning or segmentation, necessitate users to input a prompt (e.g., text or bounding boxes) to predict a unique outcome, presenting challenges such as high-cost prompt input or limited information output. This paper introduces a new task ``Image Collaborative Segmentation and Captioning'' (SegCaptioning), which aims to translate a straightforward prompt, like a bounding box around an object, into diverse semantic interpretations represented by (caption, masks) pairs, allowing flexible result selection by users. This task poses significant challenges, including accurately capturing a user's intention from a minimal prompt while simultaneously predicting multiple semantically aligned caption words and masks. Technically, we propose a novel Scene Graph Guided Diffusion Model that leverages structured scene graph features for correlated mask-caption prediction. Initially, we introduce a Prompt-Centric Scene Graph Adaptor to map a user's prompt to a scene graph, effectively capturing his intention. Subsequently, we employ a diffusion process incorporating a Scene Graph Guided Bimodal Transformer to predict correlated caption-mask pairs by uncovering intricate correlations between them. To ensure accurate alignment, we design a Multi-Entities Contrastive Learning loss to explicitly align visual and textual entities by considering inter-modal similarity, resulting in well-aligned caption-mask pairs. Extensive experiments conducted on two datasets demonstrate that SGDiff achieves superior performance in SegCaptioning, yielding promising results for both captioning and segmentation tasks with minimal prompt input.
format Preprint
id arxiv_https___arxiv_org_abs_2512_01975
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SGDiff: Scene Graph Guided Diffusion Model for Image Collaborative SegCaptioning
Zhang, Xu
Yuan, Jin
Zhang, Hanwang
Zhong, Guojin
Zang, Yongsheng
Lin, Jiacheng
Li, Zhiyong
Computer Vision and Pattern Recognition
Controllable image semantic understanding tasks, such as captioning or segmentation, necessitate users to input a prompt (e.g., text or bounding boxes) to predict a unique outcome, presenting challenges such as high-cost prompt input or limited information output. This paper introduces a new task ``Image Collaborative Segmentation and Captioning'' (SegCaptioning), which aims to translate a straightforward prompt, like a bounding box around an object, into diverse semantic interpretations represented by (caption, masks) pairs, allowing flexible result selection by users. This task poses significant challenges, including accurately capturing a user's intention from a minimal prompt while simultaneously predicting multiple semantically aligned caption words and masks. Technically, we propose a novel Scene Graph Guided Diffusion Model that leverages structured scene graph features for correlated mask-caption prediction. Initially, we introduce a Prompt-Centric Scene Graph Adaptor to map a user's prompt to a scene graph, effectively capturing his intention. Subsequently, we employ a diffusion process incorporating a Scene Graph Guided Bimodal Transformer to predict correlated caption-mask pairs by uncovering intricate correlations between them. To ensure accurate alignment, we design a Multi-Entities Contrastive Learning loss to explicitly align visual and textual entities by considering inter-modal similarity, resulting in well-aligned caption-mask pairs. Extensive experiments conducted on two datasets demonstrate that SGDiff achieves superior performance in SegCaptioning, yielding promising results for both captioning and segmentation tasks with minimal prompt input.
title SGDiff: Scene Graph Guided Diffusion Model for Image Collaborative SegCaptioning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.01975