TransRef: Multi-Scale Reference Embedding Transformer for Reference-Guided Image Inpainting

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Taorong, Liao, Liang, Chen, Delin, Xiao, Jing, Wang, Zheng, Lin, Chia-Wen, Satoh, Shin'ichi
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910820436279296
author Liu, Taorong
Liao, Liang
Chen, Delin
Xiao, Jing
Wang, Zheng
Lin, Chia-Wen
Satoh, Shin'ichi
author_facet Liu, Taorong
Liao, Liang
Chen, Delin
Xiao, Jing
Wang, Zheng
Lin, Chia-Wen
Satoh, Shin'ichi
contents Image inpainting for completing complicated semantic environments and diverse hole patterns of corrupted images is challenging even for state-of-the-art learning-based inpainting methods trained on large-scale data. A reference image capturing the same scene of a corrupted image offers informative guidance for completing the corrupted image as it shares similar texture and structure priors to that of the holes of the corrupted image. In this work, we propose a transformer-based encoder-decoder network, named TransRef, for reference-guided image inpainting. Specifically, the guidance is conducted progressively through a reference embedding procedure, in which the referencing features are subsequently aligned and fused with the features of the corrupted image. For precise utilization of the reference features for guidance, a reference-patch alignment (Ref-PA) module is proposed to align the patch features of the reference and corrupted images and harmonize their style differences, while a reference-patch transformer (Ref-PT) module is proposed to refine the embedded reference feature. Moreover, to facilitate the research of reference-guided image restoration tasks, we construct a publicly accessible benchmark dataset containing 50K pairs of input and reference images. Both quantitative and qualitative evaluations demonstrate the efficacy of the reference information and the proposed method over the state-of-the-art methods in completing complex holes. Code and dataset can be accessed at https://github.com/Cameltr/TransRef.
format Preprint
id arxiv_https___arxiv_org_abs_2306_11528
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle TransRef: Multi-Scale Reference Embedding Transformer for Reference-Guided Image Inpainting
Liu, Taorong
Liao, Liang
Chen, Delin
Xiao, Jing
Wang, Zheng
Lin, Chia-Wen
Satoh, Shin'ichi
Computer Vision and Pattern Recognition
Image inpainting for completing complicated semantic environments and diverse hole patterns of corrupted images is challenging even for state-of-the-art learning-based inpainting methods trained on large-scale data. A reference image capturing the same scene of a corrupted image offers informative guidance for completing the corrupted image as it shares similar texture and structure priors to that of the holes of the corrupted image. In this work, we propose a transformer-based encoder-decoder network, named TransRef, for reference-guided image inpainting. Specifically, the guidance is conducted progressively through a reference embedding procedure, in which the referencing features are subsequently aligned and fused with the features of the corrupted image. For precise utilization of the reference features for guidance, a reference-patch alignment (Ref-PA) module is proposed to align the patch features of the reference and corrupted images and harmonize their style differences, while a reference-patch transformer (Ref-PT) module is proposed to refine the embedded reference feature. Moreover, to facilitate the research of reference-guided image restoration tasks, we construct a publicly accessible benchmark dataset containing 50K pairs of input and reference images. Both quantitative and qualitative evaluations demonstrate the efficacy of the reference information and the proposed method over the state-of-the-art methods in completing complex holes. Code and dataset can be accessed at https://github.com/Cameltr/TransRef.
title TransRef: Multi-Scale Reference Embedding Transformer for Reference-Guided Image Inpainting
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2306.11528