Single Document Image Highlight Removal via A Large-Scale Real-World Dataset and A Location-Aware Network

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Pan, Lu, Huang, Yu-Hsuan, Xie, Hongxia, Zhang, Cheng, Zhao, Hongwei, Shuai, Hong-Han, Cheng, Wen-Huang
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916698558300160
author Pan, Lu
Huang, Yu-Hsuan
Xie, Hongxia
Zhang, Cheng
Zhao, Hongwei
Shuai, Hong-Han
Cheng, Wen-Huang
author_facet Pan, Lu
Huang, Yu-Hsuan
Xie, Hongxia
Zhang, Cheng
Zhao, Hongwei
Shuai, Hong-Han
Cheng, Wen-Huang
contents Reflective documents often suffer from specular highlights under ambient lighting, severely hindering text readability and degrading overall visual quality. Although recent deep learning methods show promise in highlight removal, they remain suboptimal for document images, primarily due to the lack of dedicated datasets and tailored architectural designs. To tackle these challenges, we present DocHR14K, a large-scale real-world dataset comprising 14,902 high-resolution image pairs across six document categories and various lighting conditions. To the best of our knowledge, this is the first high-resolution dataset for document highlight removal that captures a wide range of real-world lighting conditions. Additionally, motivated by the observation that the residual map between highlighted and clean images naturally reveals the spatial structure of highlight regions, we propose a simple yet effective Highlight Location Prior (HLP) to estimate highlight masks without human annotations. Building on this prior, we present the Location-Aware Laplacian Pyramid Highlight Removal Network (L2HRNet), which effectively removes highlights by leveraging estimated priors and incorporates diffusion module to restore details. Extensive experiments demonstrate that DocHR14K improves highlight removal under diverse lighting conditions. Our L2HRNet achieves state-of-the-art performance across three benchmark datasets, including a 5.01\% increase in PSNR and a 13.17\% reduction in RMSE on DocHR14K.
format Preprint
id arxiv_https___arxiv_org_abs_2504_14238
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Single Document Image Highlight Removal via A Large-Scale Real-World Dataset and A Location-Aware Network
Pan, Lu
Huang, Yu-Hsuan
Xie, Hongxia
Zhang, Cheng
Zhao, Hongwei
Shuai, Hong-Han
Cheng, Wen-Huang
Computer Vision and Pattern Recognition
Reflective documents often suffer from specular highlights under ambient lighting, severely hindering text readability and degrading overall visual quality. Although recent deep learning methods show promise in highlight removal, they remain suboptimal for document images, primarily due to the lack of dedicated datasets and tailored architectural designs. To tackle these challenges, we present DocHR14K, a large-scale real-world dataset comprising 14,902 high-resolution image pairs across six document categories and various lighting conditions. To the best of our knowledge, this is the first high-resolution dataset for document highlight removal that captures a wide range of real-world lighting conditions. Additionally, motivated by the observation that the residual map between highlighted and clean images naturally reveals the spatial structure of highlight regions, we propose a simple yet effective Highlight Location Prior (HLP) to estimate highlight masks without human annotations. Building on this prior, we present the Location-Aware Laplacian Pyramid Highlight Removal Network (L2HRNet), which effectively removes highlights by leveraging estimated priors and incorporates diffusion module to restore details. Extensive experiments demonstrate that DocHR14K improves highlight removal under diverse lighting conditions. Our L2HRNet achieves state-of-the-art performance across three benchmark datasets, including a 5.01\% increase in PSNR and a 13.17\% reduction in RMSE on DocHR14K.
title Single Document Image Highlight Removal via A Large-Scale Real-World Dataset and A Location-Aware Network
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2504.14238