GLM-OCR Technical Report

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Duan, Shuaiqi, Xue, Yadong, Wang, Weihan, Su, Zhe, Liu, Huan, Yang, Sheng, Gan, Guobing, Wang, Guo, Wang, Zihan, Yan, Shengdong, Jin, Dexin, Zhang, Yuxuan, Wen, Guohong, Wang, Yanfeng, Zhang, Yutao, Zhang, Xiaohan, Hong, Wenyi, Cen, Yukuo, Yin, Da, Chen, Bin, Yu, Wenmeng, Gu, Xiaotao, Tang, Jie
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917346765963264
author Duan, Shuaiqi
Xue, Yadong
Wang, Weihan
Su, Zhe
Liu, Huan
Yang, Sheng
Gan, Guobing
Wang, Guo
Wang, Zihan
Yan, Shengdong
Jin, Dexin
Zhang, Yuxuan
Wen, Guohong
Wang, Yanfeng
Zhang, Yutao
Zhang, Xiaohan
Hong, Wenyi
Cen, Yukuo
Yin, Da
Chen, Bin
Yu, Wenmeng
Gu, Xiaotao
Tang, Jie
author_facet Duan, Shuaiqi
Xue, Yadong
Wang, Weihan
Su, Zhe
Liu, Huan
Yang, Sheng
Gan, Guobing
Wang, Guo
Wang, Zihan
Yan, Shengdong
Jin, Dexin
Zhang, Yuxuan
Wen, Guohong
Wang, Yanfeng
Zhang, Yutao
Zhang, Xiaohan
Hong, Wenyi
Cen, Yukuo
Yin, Da
Chen, Bin
Yu, Wenmeng
Gu, Xiaotao
Tang, Jie
contents GLM-OCR is an efficient 0.9B-parameter compact multimodal model designed for real-world document understanding. It combines a 0.4B-parameter CogViT visual encoder with a 0.5B-parameter GLM language decoder, achieving a strong balance between computational efficiency and recognition performance. To address the inefficiency of standard autoregressive decoding in deterministic OCR tasks, GLM-OCR introduces a Multi-Token Prediction (MTP) mechanism that predicts multiple tokens per step, significantly improving decoding throughput while keeping memory overhead low through shared parameters. At the system level, a two-stage pipeline is adopted: PP-DocLayout-V3 first performs layout analysis, followed by parallel region-level recognition. Extensive evaluations on public benchmarks and industrial scenarios show that GLM-OCR achieves competitive or state-of-the-art performance in document parsing, text and formula transcription, table structure recovery, and key information extraction. Its compact architecture and structured generation make it suitable for both resource-constrained edge deployment and large-scale production systems.
format Preprint
id arxiv_https___arxiv_org_abs_2603_10910
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle GLM-OCR Technical Report
Duan, Shuaiqi
Xue, Yadong
Wang, Weihan
Su, Zhe
Liu, Huan
Yang, Sheng
Gan, Guobing
Wang, Guo
Wang, Zihan
Yan, Shengdong
Jin, Dexin
Zhang, Yuxuan
Wen, Guohong
Wang, Yanfeng
Zhang, Yutao
Zhang, Xiaohan
Hong, Wenyi
Cen, Yukuo
Yin, Da
Chen, Bin
Yu, Wenmeng
Gu, Xiaotao
Tang, Jie
Computation and Language
GLM-OCR is an efficient 0.9B-parameter compact multimodal model designed for real-world document understanding. It combines a 0.4B-parameter CogViT visual encoder with a 0.5B-parameter GLM language decoder, achieving a strong balance between computational efficiency and recognition performance. To address the inefficiency of standard autoregressive decoding in deterministic OCR tasks, GLM-OCR introduces a Multi-Token Prediction (MTP) mechanism that predicts multiple tokens per step, significantly improving decoding throughput while keeping memory overhead low through shared parameters. At the system level, a two-stage pipeline is adopted: PP-DocLayout-V3 first performs layout analysis, followed by parallel region-level recognition. Extensive evaluations on public benchmarks and industrial scenarios show that GLM-OCR achieves competitive or state-of-the-art performance in document parsing, text and formula transcription, table structure recovery, and key information extraction. Its compact architecture and structured generation make it suitable for both resource-constrained edge deployment and large-scale production systems.
title GLM-OCR Technical Report
topic Computation and Language
url https://arxiv.org/abs/2603.10910