HiFi-Inpaint: Towards High-Fidelity Reference-Based Inpainting for Generating Detail-Preserving Human-Product Images

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Yichen, Zhou, Donghao, Wang, Jie, Gao, Xin, Liu, Guisheng, Li, Jiatong, Zhang, Quanwei, Lyu, Qiang, Guo, Lanqing, Wen, Shilei, Wang, Weiqiang, Heng, Pheng-Ann
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917415361708032
author Liu, Yichen
Zhou, Donghao
Wang, Jie
Gao, Xin
Liu, Guisheng
Li, Jiatong
Zhang, Quanwei
Lyu, Qiang
Guo, Lanqing
Wen, Shilei
Wang, Weiqiang
Heng, Pheng-Ann
author_facet Liu, Yichen
Zhou, Donghao
Wang, Jie
Gao, Xin
Liu, Guisheng
Li, Jiatong
Zhang, Quanwei
Lyu, Qiang
Guo, Lanqing
Wen, Shilei
Wang, Weiqiang
Heng, Pheng-Ann
contents Human-product images, which showcase the integration of humans and products, play a vital role in advertising, e-commerce, and digital marketing. The essential challenge of generating such images lies in ensuring the high-fidelity preservation of product details. Among existing paradigms, reference-based inpainting offers a targeted solution by leveraging product reference images to guide the inpainting process. However, limitations remain in three key aspects: the lack of diverse large-scale training data, the struggle of current models to focus on product detail preservation, and the inability of coarse supervision for achieving precise guidance. To address these issues, we propose HiFi-Inpaint, a novel high-fidelity reference-based inpainting framework tailored for generating human-product images. HiFi-Inpaint introduces Shared Enhancement Attention (SEA) to refine fine-grained product features and Detail-Aware Loss (DAL) to enforce precise pixel-level supervision using high-frequency maps. Additionally, we construct a new dataset, HP-Image-40K, with samples curated from self-synthesis data and processed with automatic filtering. Experimental results show that HiFi-Inpaint achieves state-of-the-art performance, delivering detail-preserving human-product images.
format Preprint
id arxiv_https___arxiv_org_abs_2603_02210
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle HiFi-Inpaint: Towards High-Fidelity Reference-Based Inpainting for Generating Detail-Preserving Human-Product Images
Liu, Yichen
Zhou, Donghao
Wang, Jie
Gao, Xin
Liu, Guisheng
Li, Jiatong
Zhang, Quanwei
Lyu, Qiang
Guo, Lanqing
Wen, Shilei
Wang, Weiqiang
Heng, Pheng-Ann
Computer Vision and Pattern Recognition
Human-product images, which showcase the integration of humans and products, play a vital role in advertising, e-commerce, and digital marketing. The essential challenge of generating such images lies in ensuring the high-fidelity preservation of product details. Among existing paradigms, reference-based inpainting offers a targeted solution by leveraging product reference images to guide the inpainting process. However, limitations remain in three key aspects: the lack of diverse large-scale training data, the struggle of current models to focus on product detail preservation, and the inability of coarse supervision for achieving precise guidance. To address these issues, we propose HiFi-Inpaint, a novel high-fidelity reference-based inpainting framework tailored for generating human-product images. HiFi-Inpaint introduces Shared Enhancement Attention (SEA) to refine fine-grained product features and Detail-Aware Loss (DAL) to enforce precise pixel-level supervision using high-frequency maps. Additionally, we construct a new dataset, HP-Image-40K, with samples curated from self-synthesis data and processed with automatic filtering. Experimental results show that HiFi-Inpaint achieves state-of-the-art performance, delivering detail-preserving human-product images.
title HiFi-Inpaint: Towards High-Fidelity Reference-Based Inpainting for Generating Detail-Preserving Human-Product Images
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.02210