TextDoctor: Unified Document Image Inpainting via Patch Pyramid Diffusion Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Lu, Wanglong, Su, Lingming, Zheng, Jingjing, de Melo, Vinícius Veloso, Shoeleh, Farzaneh, Hawkin, John, Tricco, Terrence, Zhao, Hanli, Jiang, Xianta
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916645074632704
author Lu, Wanglong
Su, Lingming
Zheng, Jingjing
de Melo, Vinícius Veloso
Shoeleh, Farzaneh
Hawkin, John
Tricco, Terrence
Zhao, Hanli
Jiang, Xianta
author_facet Lu, Wanglong
Su, Lingming
Zheng, Jingjing
de Melo, Vinícius Veloso
Shoeleh, Farzaneh
Hawkin, John
Tricco, Terrence
Zhao, Hanli
Jiang, Xianta
contents Digital versions of real-world text documents often suffer from issues like environmental corrosion of the original document, low-quality scanning, or human interference. Existing document restoration and inpainting methods typically struggle with generalizing to unseen document styles and handling high-resolution images. To address these challenges, we introduce TextDoctor, a novel unified document image inpainting method. Inspired by human reading behavior, TextDoctor restores fundamental text elements from patches and then applies diffusion models to entire document images instead of training models on specific document types. To handle varying text sizes and avoid out-of-memory issues, common in high-resolution documents, we propose using structure pyramid prediction and patch pyramid diffusion models. These techniques leverage multiscale inputs and pyramid patches to enhance the quality of inpainting both globally and locally. Extensive qualitative and quantitative experiments on seven public datasets validated that TextDoctor outperforms state-of-the-art methods in restoring various types of high-resolution document images.
format Preprint
id arxiv_https___arxiv_org_abs_2503_04021
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TextDoctor: Unified Document Image Inpainting via Patch Pyramid Diffusion Models
Lu, Wanglong
Su, Lingming
Zheng, Jingjing
de Melo, Vinícius Veloso
Shoeleh, Farzaneh
Hawkin, John
Tricco, Terrence
Zhao, Hanli
Jiang, Xianta
Computer Vision and Pattern Recognition
Artificial Intelligence
68U10
I.4.3; I.4.4; I.4.5; I.4.9
Digital versions of real-world text documents often suffer from issues like environmental corrosion of the original document, low-quality scanning, or human interference. Existing document restoration and inpainting methods typically struggle with generalizing to unseen document styles and handling high-resolution images. To address these challenges, we introduce TextDoctor, a novel unified document image inpainting method. Inspired by human reading behavior, TextDoctor restores fundamental text elements from patches and then applies diffusion models to entire document images instead of training models on specific document types. To handle varying text sizes and avoid out-of-memory issues, common in high-resolution documents, we propose using structure pyramid prediction and patch pyramid diffusion models. These techniques leverage multiscale inputs and pyramid patches to enhance the quality of inpainting both globally and locally. Extensive qualitative and quantitative experiments on seven public datasets validated that TextDoctor outperforms state-of-the-art methods in restoring various types of high-resolution document images.
title TextDoctor: Unified Document Image Inpainting via Patch Pyramid Diffusion Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
68U10
I.4.3; I.4.4; I.4.5; I.4.9
url https://arxiv.org/abs/2503.04021