GLM-OCR Technical Report
Fuente:
arXiv
Saved in:
| Main Authors: | Duan, Shuaiqi, Xue, Yadong, Wang, Weihan, Su, Zhe, Liu, Huan, Yang, Sheng, Gan, Guobing, Wang, Guo, Wang, Zihan, Yan, Shengdong, Jin, Dexin, Zhang, Yuxuan, Wen, Guohong, Wang, Yanfeng, Zhang, Yutao, Zhang, Xiaohan, Hong, Wenyi, Cen, Yukuo, Yin, Da, Chen, Bin, Yu, Wenmeng, Gu, Xiaotao, Tang, Jie |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MathGLM-Vision: Solving Mathematical Problems with Multi-Modal Large Language Model
by: Yang, Zhen, et al.
Published: (2024)
by: Yang, Zhen, et al.
Published: (2024)
GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
by: V Team, et al.
Published: (2025)
by: V Team, et al.
Published: (2025)
GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents
by: V Team, et al.
Published: (2026)
by: V Team, et al.
Published: (2026)
GLM-TTS Technical Report
by: Cui, Jiayan, et al.
Published: (2025)
by: Cui, Jiayan, et al.
Published: (2025)
ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools
by: GLM, Team, et al.
Published: (2024)
by: GLM, Team, et al.
Published: (2024)
ChatGLM-Math: Improving Math Problem-Solving in Large Language Models with a Self-Critique Pipeline
by: Xu, Yifan, et al.
Published: (2024)
by: Xu, Yifan, et al.
Published: (2024)
Grazing duration and intensity modulate vegetation dynamics in semi-arid ecosystems with seasonal succession
by: Gan, Junhong, et al.
Published: (2025)
by: Gan, Junhong, et al.
Published: (2025)
CogAgent: A Visual Language Model for GUI Agents
by: Hong, Wenyi, et al.
Published: (2023)
by: Hong, Wenyi, et al.
Published: (2023)
Transferable Adversarial Examples with Bayes Approach
by: Fan, Mingyuan, et al.
Published: (2022)
by: Fan, Mingyuan, et al.
Published: (2022)
PaddleOCR 3.0 Technical Report
by: Cui, Cheng, et al.
Published: (2025)
by: Cui, Cheng, et al.
Published: (2025)
GLM-5: from Vibe Coding to Agentic Engineering
by: GLM-5-Team, et al.
Published: (2026)
by: GLM-5-Team, et al.
Published: (2026)
HunyuanOCR Technical Report
by: Hunyuan Vision Team, et al.
Published: (2025)
by: Hunyuan Vision Team, et al.
Published: (2025)
AutoWebGLM: A Large Language Model-based Web Navigating Agent
by: Lai, Hanyu, et al.
Published: (2024)
by: Lai, Hanyu, et al.
Published: (2024)
GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models
by: 5 Team, et al.
Published: (2025)
by: 5 Team, et al.
Published: (2025)
LVBench: An Extreme Long Video Understanding Benchmark
by: Wang, Weihan, et al.
Published: (2024)
by: Wang, Weihan, et al.
Published: (2024)
Ocean-OCR: Towards General OCR Application via a Vision-Language Model
by: Chen, Song, et al.
Published: (2025)
by: Chen, Song, et al.
Published: (2025)
MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models
by: Hong, Wenyi, et al.
Published: (2025)
by: Hong, Wenyi, et al.
Published: (2025)
OmniOCR: Generalist OCR for Ethnic Minority Languages
by: Liu, Bonan, et al.
Published: (2026)
by: Liu, Bonan, et al.
Published: (2026)
UI2Code^N: UI-to-Code Generation as Interactive Visual Optimization
by: Yang, Zhen, et al.
Published: (2025)
by: Yang, Zhen, et al.
Published: (2025)
ChatGLM-RLHF: Practices of Aligning Large Language Models with Human Feedback
by: Hou, Zhenyu, et al.
Published: (2024)
by: Hou, Zhenyu, et al.
Published: (2024)
Refiner: Data Refining against Gradient Leakage Attacks in Federated Learning
by: Fan, Mingyuan, et al.
Published: (2022)
by: Fan, Mingyuan, et al.
Published: (2022)
AutoGLM: Autonomous Foundation Agents for GUIs
by: Liu, Xiao, et al.
Published: (2024)
by: Liu, Xiao, et al.
Published: (2024)
PrefRAG: Preference-Driven Multi-Source Retrieval Augmented Generation
by: Zhao, Qingfei, et al.
Published: (2024)
by: Zhao, Qingfei, et al.
Published: (2024)
FeTaX2: A ferrimagnetic quantum anomalous Hall insulator
by: Jiang, Yadong, et al.
Published: (2025)
by: Jiang, Yadong, et al.
Published: (2025)
CogVLM2: Visual Language Models for Image and Video Understanding
by: Hong, Wenyi, et al.
Published: (2024)
by: Hong, Wenyi, et al.
Published: (2024)
Effect of protection zone on the dynamics of a diffusion-advection population-toxicant model
by: Gao, Jing, et al.
Published: (2025)
by: Gao, Jing, et al.
Published: (2025)
Study on Transverse Surface Cracks in the Continuous Casting Slab of a Microalloyed Steel
by: Yadong Wang, et al.
Published: (2025)
by: Yadong Wang, et al.
Published: (2025)
A Review of Macrosegregation Simulation for the Continuous Casting Process
by: Yadong Wang, et al.
Published: (2025)
by: Yadong Wang, et al.
Published: (2025)
PST-Bench: Tracing and Benchmarking the Source of Publications
by: Zhang, Fanjin, et al.
Published: (2024)
by: Zhang, Fanjin, et al.
Published: (2024)
PlotGen-Bench: Evaluating VLMs on Generating Visualization Code from Diverse Plots across Multiple Libraries
by: Zhao, Yi, et al.
Published: (2025)
by: Zhao, Yi, et al.
Published: (2025)
Glyph: Scaling Context Windows via Visual-Text Compression
by: Cheng, Jiale, et al.
Published: (2025)
by: Cheng, Jiale, et al.
Published: (2025)
DPSW-Sketch: A Differentially Private Sketch Framework for Frequency Estimation over Sliding Windows (Technical Report)
by: Wang, Yiping, et al.
Published: (2024)
by: Wang, Yiping, et al.
Published: (2024)
Antibiotic‐Mediated Microbiota Depletion Suggests an Association Between Gastric Juice Dysbacteriosis and Abnormal Bile Acid Metabolism in Chronic Atrophic Gastritis Rats
by: Yan Zhang, et al.
Published: (2026)
by: Yan Zhang, et al.
Published: (2026)
HCMRM: A High-Consistency Multimodal Relevance Model for Search Ads
by: Gan, Guobing, et al.
Published: (2025)
by: Gan, Guobing, et al.
Published: (2025)
M$^3$AV: A Multimodal, Multigenre, and Multipurpose Audio-Visual Academic Lecture Dataset
by: Chen, Zhe, et al.
Published: (2024)
by: Chen, Zhe, et al.
Published: (2024)
CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer
by: Yang, Zhuoyi, et al.
Published: (2024)
by: Yang, Zhuoyi, et al.
Published: (2024)
When Good OCR Is Not Enough: Benchmarking OCR Robustness for Retrieval-Augmented Generation
by: Sun, Lin, et al.
Published: (2026)
by: Sun, Lin, et al.
Published: (2026)
Recent Progress of Corrosion Prevention Method of Magnesium Alloy
by: Qi He, et al.
Published: (2024)
by: Qi He, et al.
Published: (2024)
LongRAG: A Dual-Perspective Retrieval-Augmented Generation Paradigm for Long-Context Question Answering
by: Zhao, Qingfei, et al.
Published: (2024)
by: Zhao, Qingfei, et al.
Published: (2024)
MME-VideoOCR: Evaluating OCR-Based Capabilities of Multimodal LLMs in Video Scenarios
by: Shi, Yang, et al.
Published: (2025)
by: Shi, Yang, et al.
Published: (2025)
Similar Items
-
MathGLM-Vision: Solving Mathematical Problems with Multi-Modal Large Language Model
by: Yang, Zhen, et al.
Published: (2024) -
GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
by: V Team, et al.
Published: (2025) -
GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents
by: V Team, et al.
Published: (2026) -
GLM-TTS Technical Report
by: Cui, Jiayan, et al.
Published: (2025) -
ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools
by: GLM, Team, et al.
Published: (2024)