GLM-OCR Technical Report
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917346765963264 |
|---|---|
| author | Duan, Shuaiqi Xue, Yadong Wang, Weihan Su, Zhe Liu, Huan Yang, Sheng Gan, Guobing Wang, Guo Wang, Zihan Yan, Shengdong Jin, Dexin Zhang, Yuxuan Wen, Guohong Wang, Yanfeng Zhang, Yutao Zhang, Xiaohan Hong, Wenyi Cen, Yukuo Yin, Da Chen, Bin Yu, Wenmeng Gu, Xiaotao Tang, Jie |
| author_facet | Duan, Shuaiqi Xue, Yadong Wang, Weihan Su, Zhe Liu, Huan Yang, Sheng Gan, Guobing Wang, Guo Wang, Zihan Yan, Shengdong Jin, Dexin Zhang, Yuxuan Wen, Guohong Wang, Yanfeng Zhang, Yutao Zhang, Xiaohan Hong, Wenyi Cen, Yukuo Yin, Da Chen, Bin Yu, Wenmeng Gu, Xiaotao Tang, Jie |
| contents | GLM-OCR is an efficient 0.9B-parameter compact multimodal model designed for real-world document understanding. It combines a 0.4B-parameter CogViT visual encoder with a 0.5B-parameter GLM language decoder, achieving a strong balance between computational efficiency and recognition performance. To address the inefficiency of standard autoregressive decoding in deterministic OCR tasks, GLM-OCR introduces a Multi-Token Prediction (MTP) mechanism that predicts multiple tokens per step, significantly improving decoding throughput while keeping memory overhead low through shared parameters. At the system level, a two-stage pipeline is adopted: PP-DocLayout-V3 first performs layout analysis, followed by parallel region-level recognition. Extensive evaluations on public benchmarks and industrial scenarios show that GLM-OCR achieves competitive or state-of-the-art performance in document parsing, text and formula transcription, table structure recovery, and key information extraction. Its compact architecture and structured generation make it suitable for both resource-constrained edge deployment and large-scale production systems. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2603_10910 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | GLM-OCR Technical Report Duan, Shuaiqi Xue, Yadong Wang, Weihan Su, Zhe Liu, Huan Yang, Sheng Gan, Guobing Wang, Guo Wang, Zihan Yan, Shengdong Jin, Dexin Zhang, Yuxuan Wen, Guohong Wang, Yanfeng Zhang, Yutao Zhang, Xiaohan Hong, Wenyi Cen, Yukuo Yin, Da Chen, Bin Yu, Wenmeng Gu, Xiaotao Tang, Jie Computation and Language GLM-OCR is an efficient 0.9B-parameter compact multimodal model designed for real-world document understanding. It combines a 0.4B-parameter CogViT visual encoder with a 0.5B-parameter GLM language decoder, achieving a strong balance between computational efficiency and recognition performance. To address the inefficiency of standard autoregressive decoding in deterministic OCR tasks, GLM-OCR introduces a Multi-Token Prediction (MTP) mechanism that predicts multiple tokens per step, significantly improving decoding throughput while keeping memory overhead low through shared parameters. At the system level, a two-stage pipeline is adopted: PP-DocLayout-V3 first performs layout analysis, followed by parallel region-level recognition. Extensive evaluations on public benchmarks and industrial scenarios show that GLM-OCR achieves competitive or state-of-the-art performance in document parsing, text and formula transcription, table structure recovery, and key information extraction. Its compact architecture and structured generation make it suitable for both resource-constrained edge deployment and large-scale production systems. |
| title | GLM-OCR Technical Report |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2603.10910 |