Platypus: A Generalized Specialist Model for Reading Text in Various Forms

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Peng, Li, Zhaohai, Tang, Jun, Zhong, Humen, Huang, Fei, Yang, Zhibo, Yao, Cong
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917759761252352
author Wang, Peng
Li, Zhaohai
Tang, Jun
Zhong, Humen
Huang, Fei
Yang, Zhibo
Yao, Cong
author_facet Wang, Peng
Li, Zhaohai
Tang, Jun
Zhong, Humen
Huang, Fei
Yang, Zhibo
Yao, Cong
contents Reading text from images (either natural scenes or documents) has been a long-standing research topic for decades, due to the high technical challenge and wide application range. Previously, individual specialist models are developed to tackle the sub-tasks of text reading (e.g., scene text recognition, handwritten text recognition and mathematical expression recognition). However, such specialist models usually cannot effectively generalize across different sub-tasks. Recently, generalist models (such as GPT-4V), trained on tremendous data in a unified way, have shown enormous potential in reading text in various scenarios, but with the drawbacks of limited accuracy and low efficiency. In this work, we propose Platypus, a generalized specialist model for text reading. Specifically, Platypus combines the best of both worlds: being able to recognize text of various forms with a single unified architecture, while achieving excellent accuracy and high efficiency. To better exploit the advantage of Platypus, we also construct a text reading dataset (called Worms), the images of which are curated from previous datasets and partially re-labeled. Experiments on standard benchmarks demonstrate the effectiveness and superiority of the proposed Platypus model. Model and data will be made publicly available at https://github.com/AlibabaResearch/AdvancedLiterateMachinery/tree/main/OCR/Platypus.
format Preprint
id arxiv_https___arxiv_org_abs_2408_14805
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Platypus: A Generalized Specialist Model for Reading Text in Various Forms
Wang, Peng
Li, Zhaohai
Tang, Jun
Zhong, Humen
Huang, Fei
Yang, Zhibo
Yao, Cong
Computer Vision and Pattern Recognition
Reading text from images (either natural scenes or documents) has been a long-standing research topic for decades, due to the high technical challenge and wide application range. Previously, individual specialist models are developed to tackle the sub-tasks of text reading (e.g., scene text recognition, handwritten text recognition and mathematical expression recognition). However, such specialist models usually cannot effectively generalize across different sub-tasks. Recently, generalist models (such as GPT-4V), trained on tremendous data in a unified way, have shown enormous potential in reading text in various scenarios, but with the drawbacks of limited accuracy and low efficiency. In this work, we propose Platypus, a generalized specialist model for text reading. Specifically, Platypus combines the best of both worlds: being able to recognize text of various forms with a single unified architecture, while achieving excellent accuracy and high efficiency. To better exploit the advantage of Platypus, we also construct a text reading dataset (called Worms), the images of which are curated from previous datasets and partially re-labeled. Experiments on standard benchmarks demonstrate the effectiveness and superiority of the proposed Platypus model. Model and data will be made publicly available at https://github.com/AlibabaResearch/AdvancedLiterateMachinery/tree/main/OCR/Platypus.
title Platypus: A Generalized Specialist Model for Reading Text in Various Forms
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2408.14805