UPOCR: Towards Unified Pixel-Level OCR Interface

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Peng, Dezhi, Yang, Zhenhua, Zhang, Jiaxin, Liu, Chongyu, Shi, Yongxin, Ding, Kai, Guo, Fengjun, Jin, Lianwen
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917532465627136
author Peng, Dezhi
Yang, Zhenhua
Zhang, Jiaxin
Liu, Chongyu
Shi, Yongxin
Ding, Kai
Guo, Fengjun
Jin, Lianwen
author_facet Peng, Dezhi
Yang, Zhenhua
Zhang, Jiaxin
Liu, Chongyu
Shi, Yongxin
Ding, Kai
Guo, Fengjun
Jin, Lianwen
contents Existing optical character recognition (OCR) methods rely on task-specific designs with divergent paradigms, architectures, and training strategies, which significantly increases the complexity of research and maintenance and hinders the fast deployment in applications. To this end, we propose UPOCR, a simple-yet-effective generalist model for Unified Pixel-level OCR interface. Specifically, the UPOCR unifies the paradigm of diverse OCR tasks as image-to-image transformation and the architecture as a vision Transformer (ViT)-based encoder-decoder with learnable task prompts. The prompts push the general feature representations extracted by the encoder towards task-specific spaces, endowing the decoder with task awareness. Moreover, the model training is uniformly aimed at minimizing the discrepancy between the predicted and ground-truth images regardless of the inhomogeneity among tasks. Experiments are conducted on three pixel-level OCR tasks including text removal, text segmentation, and tampered text detection. Without bells and whistles, the experimental results showcase that the proposed method can simultaneously achieve state-of-the-art performance on three tasks with a unified single model, which provides valuable strategies and insights for future research on generalist OCR models. Code is available at https://github.com/shannanyinxiang/UPOCR.
format Preprint
id arxiv_https___arxiv_org_abs_2312_02694
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle UPOCR: Towards Unified Pixel-Level OCR Interface
Peng, Dezhi
Yang, Zhenhua
Zhang, Jiaxin
Liu, Chongyu
Shi, Yongxin
Ding, Kai
Guo, Fengjun
Jin, Lianwen
Computer Vision and Pattern Recognition
Existing optical character recognition (OCR) methods rely on task-specific designs with divergent paradigms, architectures, and training strategies, which significantly increases the complexity of research and maintenance and hinders the fast deployment in applications. To this end, we propose UPOCR, a simple-yet-effective generalist model for Unified Pixel-level OCR interface. Specifically, the UPOCR unifies the paradigm of diverse OCR tasks as image-to-image transformation and the architecture as a vision Transformer (ViT)-based encoder-decoder with learnable task prompts. The prompts push the general feature representations extracted by the encoder towards task-specific spaces, endowing the decoder with task awareness. Moreover, the model training is uniformly aimed at minimizing the discrepancy between the predicted and ground-truth images regardless of the inhomogeneity among tasks. Experiments are conducted on three pixel-level OCR tasks including text removal, text segmentation, and tampered text detection. Without bells and whistles, the experimental results showcase that the proposed method can simultaneously achieve state-of-the-art performance on three tasks with a unified single model, which provides valuable strategies and insights for future research on generalist OCR models. Code is available at https://github.com/shannanyinxiang/UPOCR.
title UPOCR: Towards Unified Pixel-Level OCR Interface
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2312.02694