Zippo: Zipping Color and Transparency Distributions into a Single Diffusion Model
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866913271452270592 |
|---|---|
| author | Xie, Kangyang Yang, Binbin Chen, Hao Wang, Meng Zou, Cheng Xue, Hui Yang, Ming Shen, Chunhua |
| author_facet | Xie, Kangyang Yang, Binbin Chen, Hao Wang, Meng Zou, Cheng Xue, Hui Yang, Ming Shen, Chunhua |
| contents | Beyond the superiority of the text-to-image diffusion model in generating high-quality images, recent studies have attempted to uncover its potential for adapting the learned semantic knowledge to visual perception tasks. In this work, instead of translating a generative diffusion model into a visual perception model, we explore to retain the generative ability with the perceptive adaptation. To accomplish this, we present Zippo, a unified framework for zipping the color and transparency distributions into a single diffusion model by expanding the diffusion latent into a joint representation of RGB images and alpha mattes. By alternatively selecting one modality as the condition and then applying the diffusion process to the counterpart modality, Zippo is capable of generating RGB images from alpha mattes and predicting transparency from input images. In addition to single-modality prediction, we propose a modality-aware noise reassignment strategy to further empower Zippo with jointly generating RGB images and its corresponding alpha mattes under the text guidance. Our experiments showcase Zippo's ability of efficient text-conditioned transparent image generation and present plausible results of Matte-to-RGB and RGB-to-Matte translation. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2403_11077 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Zippo: Zipping Color and Transparency Distributions into a Single Diffusion Model Xie, Kangyang Yang, Binbin Chen, Hao Wang, Meng Zou, Cheng Xue, Hui Yang, Ming Shen, Chunhua Computer Vision and Pattern Recognition Beyond the superiority of the text-to-image diffusion model in generating high-quality images, recent studies have attempted to uncover its potential for adapting the learned semantic knowledge to visual perception tasks. In this work, instead of translating a generative diffusion model into a visual perception model, we explore to retain the generative ability with the perceptive adaptation. To accomplish this, we present Zippo, a unified framework for zipping the color and transparency distributions into a single diffusion model by expanding the diffusion latent into a joint representation of RGB images and alpha mattes. By alternatively selecting one modality as the condition and then applying the diffusion process to the counterpart modality, Zippo is capable of generating RGB images from alpha mattes and predicting transparency from input images. In addition to single-modality prediction, we propose a modality-aware noise reassignment strategy to further empower Zippo with jointly generating RGB images and its corresponding alpha mattes under the text guidance. Our experiments showcase Zippo's ability of efficient text-conditioned transparent image generation and present plausible results of Matte-to-RGB and RGB-to-Matte translation. |
| title | Zippo: Zipping Color and Transparency Distributions into a Single Diffusion Model |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2403.11077 |