Zippo: Zipping Color and Transparency Distributions into a Single Diffusion Model

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Xie, Kangyang, Yang, Binbin, Chen, Hao, Wang, Meng, Zou, Cheng, Xue, Hui, Yang, Ming, Shen, Chunhua
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913271452270592
author Xie, Kangyang
Yang, Binbin
Chen, Hao
Wang, Meng
Zou, Cheng
Xue, Hui
Yang, Ming
Shen, Chunhua
author_facet Xie, Kangyang
Yang, Binbin
Chen, Hao
Wang, Meng
Zou, Cheng
Xue, Hui
Yang, Ming
Shen, Chunhua
contents Beyond the superiority of the text-to-image diffusion model in generating high-quality images, recent studies have attempted to uncover its potential for adapting the learned semantic knowledge to visual perception tasks. In this work, instead of translating a generative diffusion model into a visual perception model, we explore to retain the generative ability with the perceptive adaptation. To accomplish this, we present Zippo, a unified framework for zipping the color and transparency distributions into a single diffusion model by expanding the diffusion latent into a joint representation of RGB images and alpha mattes. By alternatively selecting one modality as the condition and then applying the diffusion process to the counterpart modality, Zippo is capable of generating RGB images from alpha mattes and predicting transparency from input images. In addition to single-modality prediction, we propose a modality-aware noise reassignment strategy to further empower Zippo with jointly generating RGB images and its corresponding alpha mattes under the text guidance. Our experiments showcase Zippo's ability of efficient text-conditioned transparent image generation and present plausible results of Matte-to-RGB and RGB-to-Matte translation.
format Preprint
id arxiv_https___arxiv_org_abs_2403_11077
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Zippo: Zipping Color and Transparency Distributions into a Single Diffusion Model
Xie, Kangyang
Yang, Binbin
Chen, Hao
Wang, Meng
Zou, Cheng
Xue, Hui
Yang, Ming
Shen, Chunhua
Computer Vision and Pattern Recognition
Beyond the superiority of the text-to-image diffusion model in generating high-quality images, recent studies have attempted to uncover its potential for adapting the learned semantic knowledge to visual perception tasks. In this work, instead of translating a generative diffusion model into a visual perception model, we explore to retain the generative ability with the perceptive adaptation. To accomplish this, we present Zippo, a unified framework for zipping the color and transparency distributions into a single diffusion model by expanding the diffusion latent into a joint representation of RGB images and alpha mattes. By alternatively selecting one modality as the condition and then applying the diffusion process to the counterpart modality, Zippo is capable of generating RGB images from alpha mattes and predicting transparency from input images. In addition to single-modality prediction, we propose a modality-aware noise reassignment strategy to further empower Zippo with jointly generating RGB images and its corresponding alpha mattes under the text guidance. Our experiments showcase Zippo's ability of efficient text-conditioned transparent image generation and present plausible results of Matte-to-RGB and RGB-to-Matte translation.
title Zippo: Zipping Color and Transparency Distributions into a Single Diffusion Model
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2403.11077