Ultra High-Resolution Image Inpainting with Patch-Based Content Consistency Adapter

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhang, Jianhui, Cheng, Sheng, Sun, Qirui, Liu, Jia, Luyang, Wang, Feng, Chaoyu, Fang, Chen, Lei, Lei, Wang, Jue, Liu, Shuaicheng
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918161007247360
author Zhang, Jianhui
Cheng, Sheng
Sun, Qirui
Liu, Jia
Luyang, Wang
Feng, Chaoyu
Fang, Chen
Lei, Lei
Wang, Jue
Liu, Shuaicheng
author_facet Zhang, Jianhui
Cheng, Sheng
Sun, Qirui
Liu, Jia
Luyang, Wang
Feng, Chaoyu
Fang, Chen
Lei, Lei
Wang, Jue
Liu, Shuaicheng
contents In this work, we present Patch-Adapter, an effective framework for high-resolution text-guided image inpainting. Unlike existing methods limited to lower resolutions, our approach achieves 4K+ resolution while maintaining precise content consistency and prompt alignment, two critical challenges in image inpainting that intensify with increasing resolution and texture complexity. Patch-Adapter leverages a two-stage adapter architecture to scale the diffusion model's resolution from 1K to 4K+ without requiring structural overhauls: (1) Dual Context Adapter learns coherence between masked and unmasked regions at reduced resolutions to establish global structural consistency; and (2) Reference Patch Adapter implements a patch-level attention mechanism for full-resolution inpainting, preserving local detail fidelity through adaptive feature fusion. This dual-stage architecture uniquely addresses the scalability gap in high-resolution inpainting by decoupling global semantics from localized refinement. Experiments demonstrate that Patch-Adapter not only resolves artifacts common in large-scale inpainting but also achieves state-of-the-art performance on the OpenImages and Photo-Concept-Bucket datasets, outperforming existing methods in both perceptual quality and text-prompt adherence.
format Preprint
id arxiv_https___arxiv_org_abs_2510_13419
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Ultra High-Resolution Image Inpainting with Patch-Based Content Consistency Adapter
Zhang, Jianhui
Cheng, Sheng
Sun, Qirui
Liu, Jia
Luyang, Wang
Feng, Chaoyu
Fang, Chen
Lei, Lei
Wang, Jue
Liu, Shuaicheng
Computer Vision and Pattern Recognition
In this work, we present Patch-Adapter, an effective framework for high-resolution text-guided image inpainting. Unlike existing methods limited to lower resolutions, our approach achieves 4K+ resolution while maintaining precise content consistency and prompt alignment, two critical challenges in image inpainting that intensify with increasing resolution and texture complexity. Patch-Adapter leverages a two-stage adapter architecture to scale the diffusion model's resolution from 1K to 4K+ without requiring structural overhauls: (1) Dual Context Adapter learns coherence between masked and unmasked regions at reduced resolutions to establish global structural consistency; and (2) Reference Patch Adapter implements a patch-level attention mechanism for full-resolution inpainting, preserving local detail fidelity through adaptive feature fusion. This dual-stage architecture uniquely addresses the scalability gap in high-resolution inpainting by decoupling global semantics from localized refinement. Experiments demonstrate that Patch-Adapter not only resolves artifacts common in large-scale inpainting but also achieves state-of-the-art performance on the OpenImages and Photo-Concept-Bucket datasets, outperforming existing methods in both perceptual quality and text-prompt adherence.
title Ultra High-Resolution Image Inpainting with Patch-Based Content Consistency Adapter
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.13419