P3S-Diffusion:A Selective Subject-driven Generation Framework via Point Supervision

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hu, Junjie, Gao, Shuyong, Hong, Lingyi, Wang, Qishan, Zhao, Yuzhou, Wang, Yan, Zhang, Wenqiang
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912177465589760
author Hu, Junjie
Gao, Shuyong
Hong, Lingyi
Wang, Qishan
Zhao, Yuzhou
Wang, Yan
Zhang, Wenqiang
author_facet Hu, Junjie
Gao, Shuyong
Hong, Lingyi
Wang, Qishan
Zhao, Yuzhou
Wang, Yan
Zhang, Wenqiang
contents Recent research in subject-driven generation increasingly emphasizes the importance of selective subject features. Nevertheless, accurately selecting the content in a given reference image still poses challenges, especially when selecting the similar subjects in an image (e.g., two different dogs). Some methods attempt to use text prompts or pixel masks to isolate specific elements. However, text prompts often fall short in precisely describing specific content, and pixel masks are often expensive. To address this, we introduce P3S-Diffusion, a novel architecture designed for context-selected subject-driven generation via point supervision. P3S-Diffusion leverages minimal cost label (e.g., points) to generate subject-driven images. During fine-tuning, it can generate an expanded base mask from these points, obviating the need for additional segmentation models. The mask is employed for inpainting and aligning with subject representation. The P3S-Diffusion preserves fine features of the subjects through Multi-layers Condition Injection. Enhanced by the Attention Consistency Loss for improved training, extensive experiments demonstrate its excellent feature preservation and image generation capabilities.
format Preprint
id arxiv_https___arxiv_org_abs_2412_19533
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle P3S-Diffusion:A Selective Subject-driven Generation Framework via Point Supervision
Hu, Junjie
Gao, Shuyong
Hong, Lingyi
Wang, Qishan
Zhao, Yuzhou
Wang, Yan
Zhang, Wenqiang
Computer Vision and Pattern Recognition
Artificial Intelligence
Recent research in subject-driven generation increasingly emphasizes the importance of selective subject features. Nevertheless, accurately selecting the content in a given reference image still poses challenges, especially when selecting the similar subjects in an image (e.g., two different dogs). Some methods attempt to use text prompts or pixel masks to isolate specific elements. However, text prompts often fall short in precisely describing specific content, and pixel masks are often expensive. To address this, we introduce P3S-Diffusion, a novel architecture designed for context-selected subject-driven generation via point supervision. P3S-Diffusion leverages minimal cost label (e.g., points) to generate subject-driven images. During fine-tuning, it can generate an expanded base mask from these points, obviating the need for additional segmentation models. The mask is employed for inpainting and aligning with subject representation. The P3S-Diffusion preserves fine features of the subjects through Multi-layers Condition Injection. Enhanced by the Attention Consistency Loss for improved training, extensive experiments demonstrate its excellent feature preservation and image generation capabilities.
title P3S-Diffusion:A Selective Subject-driven Generation Framework via Point Supervision
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2412.19533