KOSMOS-2.5: A Multimodal Literate Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lv, Tengchao, Huang, Yupan, Chen, Jingye, Zhao, Yuzhong, Jia, Yilin, Cui, Lei, Ma, Shuming, Chang, Yaoyao, Huang, Shaohan, Wang, Wenhui, Dong, Li, Luo, Weiyao, Wu, Shaoxiang, Wang, Guoxin, Zhang, Cha, Wei, Furu
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913474840363008
author Lv, Tengchao
Huang, Yupan
Chen, Jingye
Zhao, Yuzhong
Jia, Yilin
Cui, Lei
Ma, Shuming
Chang, Yaoyao
Huang, Shaohan
Wang, Wenhui
Dong, Li
Luo, Weiyao
Wu, Shaoxiang
Wang, Guoxin
Zhang, Cha
Wei, Furu
author_facet Lv, Tengchao
Huang, Yupan
Chen, Jingye
Zhao, Yuzhong
Jia, Yilin
Cui, Lei
Ma, Shuming
Chang, Yaoyao
Huang, Shaohan
Wang, Wenhui
Dong, Li
Luo, Weiyao
Wu, Shaoxiang
Wang, Guoxin
Zhang, Cha
Wei, Furu
contents The automatic reading of text-intensive images represents a significant advancement toward achieving Artificial General Intelligence (AGI). In this paper we present KOSMOS-2.5, a multimodal literate model for machine reading of text-intensive images. Pre-trained on a large-scale corpus of text-intensive images, KOSMOS-2.5 excels in two distinct yet complementary transcription tasks: (1) generating spatially-aware text blocks, where each block of text is assigned spatial coordinates within the image, and (2) producing structured text output that captures both style and structure in markdown format. This unified multimodal literate capability is achieved through a shared decoder-only autoregressive Transformer architecture and task-specific prompts. Building on this foundation, we fine-tune KOSMOS-2.5 for document understanding tasks, resulting in a document understanding generalist named KOSMOS-2.5-CHAT. Additionally, a large corpus of 357.4 million document pages spanning diverse domains was curated for pre-training. We evaluate KOSMOS-2.5 on two newly proposed benchmarks, OCREval and MarkdownEval, for document-level text recognition and image-to-markdown generation, demonstrating impressive literate capabilities comparable to GPT-4o. KOSMOS-2.5-CHAT achieves performance comparable to other state-of-the-art generalists that are five times larger (1.3B vs. 7B) across nine text-rich visual question answering benchmarks. Models and code have been available at \url{https://aka.ms/kosmos25}.
format Preprint
id arxiv_https___arxiv_org_abs_2309_11419
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle KOSMOS-2.5: A Multimodal Literate Model
Lv, Tengchao
Huang, Yupan
Chen, Jingye
Zhao, Yuzhong
Jia, Yilin
Cui, Lei
Ma, Shuming
Chang, Yaoyao
Huang, Shaohan
Wang, Wenhui
Dong, Li
Luo, Weiyao
Wu, Shaoxiang
Wang, Guoxin
Zhang, Cha
Wei, Furu
Computation and Language
Computer Vision and Pattern Recognition
The automatic reading of text-intensive images represents a significant advancement toward achieving Artificial General Intelligence (AGI). In this paper we present KOSMOS-2.5, a multimodal literate model for machine reading of text-intensive images. Pre-trained on a large-scale corpus of text-intensive images, KOSMOS-2.5 excels in two distinct yet complementary transcription tasks: (1) generating spatially-aware text blocks, where each block of text is assigned spatial coordinates within the image, and (2) producing structured text output that captures both style and structure in markdown format. This unified multimodal literate capability is achieved through a shared decoder-only autoregressive Transformer architecture and task-specific prompts. Building on this foundation, we fine-tune KOSMOS-2.5 for document understanding tasks, resulting in a document understanding generalist named KOSMOS-2.5-CHAT. Additionally, a large corpus of 357.4 million document pages spanning diverse domains was curated for pre-training. We evaluate KOSMOS-2.5 on two newly proposed benchmarks, OCREval and MarkdownEval, for document-level text recognition and image-to-markdown generation, demonstrating impressive literate capabilities comparable to GPT-4o. KOSMOS-2.5-CHAT achieves performance comparable to other state-of-the-art generalists that are five times larger (1.3B vs. 7B) across nine text-rich visual question answering benchmarks. Models and code have been available at \url{https://aka.ms/kosmos25}.
title KOSMOS-2.5: A Multimodal Literate Model
topic Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2309.11419