FashionR2R: Texture-preserving Rendered-to-Real Image Translation with Diffusion Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hu, Rui, He, Qian, He, Gaofeng, Zhuang, Jiedong, Chen, Huang, Liu, Huafeng, Wang, Huamin
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929550199357440
author Hu, Rui
He, Qian
He, Gaofeng
Zhuang, Jiedong
Chen, Huang
Liu, Huafeng
Wang, Huamin
author_facet Hu, Rui
He, Qian
He, Gaofeng
Zhuang, Jiedong
Chen, Huang
Liu, Huafeng
Wang, Huamin
contents Modeling and producing lifelike clothed human images has attracted researchers' attention from different areas for decades, with the complexity from highly articulated and structured content. Rendering algorithms decompose and simulate the imaging process of a camera, while are limited by the accuracy of modeled variables and the efficiency of computation. Generative models can produce impressively vivid human images, however still lacking in controllability and editability. This paper studies photorealism enhancement of rendered images, leveraging generative power from diffusion models on the controlled basis of rendering. We introduce a novel framework to translate rendered images into their realistic counterparts, which consists of two stages: Domain Knowledge Injection (DKI) and Realistic Image Generation (RIG). In DKI, we adopt positive (real) domain finetuning and negative (rendered) domain embedding to inject knowledge into a pretrained Text-to-image (T2I) diffusion model. In RIG, we generate the realistic image corresponding to the input rendered image, with a Texture-preserving Attention Control (TAC) to preserve fine-grained clothing textures, exploiting the decoupled features encoded in the UNet structure. Additionally, we introduce SynFashion dataset, featuring high-quality digital clothing images with diverse textures. Extensive experimental results demonstrate the superiority and effectiveness of our method in rendered-to-real image translation.
format Preprint
id arxiv_https___arxiv_org_abs_2410_14429
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle FashionR2R: Texture-preserving Rendered-to-Real Image Translation with Diffusion Models
Hu, Rui
He, Qian
He, Gaofeng
Zhuang, Jiedong
Chen, Huang
Liu, Huafeng
Wang, Huamin
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Modeling and producing lifelike clothed human images has attracted researchers' attention from different areas for decades, with the complexity from highly articulated and structured content. Rendering algorithms decompose and simulate the imaging process of a camera, while are limited by the accuracy of modeled variables and the efficiency of computation. Generative models can produce impressively vivid human images, however still lacking in controllability and editability. This paper studies photorealism enhancement of rendered images, leveraging generative power from diffusion models on the controlled basis of rendering. We introduce a novel framework to translate rendered images into their realistic counterparts, which consists of two stages: Domain Knowledge Injection (DKI) and Realistic Image Generation (RIG). In DKI, we adopt positive (real) domain finetuning and negative (rendered) domain embedding to inject knowledge into a pretrained Text-to-image (T2I) diffusion model. In RIG, we generate the realistic image corresponding to the input rendered image, with a Texture-preserving Attention Control (TAC) to preserve fine-grained clothing textures, exploiting the decoupled features encoded in the UNet structure. Additionally, we introduce SynFashion dataset, featuring high-quality digital clothing images with diverse textures. Extensive experimental results demonstrate the superiority and effectiveness of our method in rendered-to-real image translation.
title FashionR2R: Texture-preserving Rendered-to-Real Image Translation with Diffusion Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2410.14429