OCRVerse: Towards Holistic OCR in End-to-End Vision-Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Zhong, Yufeng, Chen, Lei, Zhao, Xuanle, Han, Wenkang, Zheng, Liming, Huang, Jing, Jiang, Deyang, Cao, Yilin, Ma, Lin, Zeng, Zhixiong |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Reading or Reasoning? Format Decoupled Reinforcement Learning for Document OCR
by: Zhong, Yufeng, et al.
Published: (2025)
by: Zhong, Yufeng, et al.
Published: (2025)
VinciCoder: Unifying Multimodal Code Generation via Coarse-to-fine Visual Reinforcement Learning
by: Zhao, Xuanle, et al.
Published: (2025)
by: Zhao, Xuanle, et al.
Published: (2025)
Breaking the SFT Plateau: Multimodal Structured Reinforcement Learning for Chart-to-Code Generation
by: Chen, Lei, et al.
Published: (2025)
by: Chen, Lei, et al.
Published: (2025)
Chart-R1: Chain-of-Thought Supervision and Reinforcement for Advanced Chart Reasoner
by: Chen, Lei, et al.
Published: (2025)
by: Chen, Lei, et al.
Published: (2025)
ScaleTrack: Scaling and back-tracking Automated GUI Agents
by: Huang, Jing, et al.
Published: (2025)
by: Huang, Jing, et al.
Published: (2025)
TreeCUA: Efficiently Scaling GUI Automation with Tree-Structured Verifiable Evolution
by: Jiang, Deyang, et al.
Published: (2026)
by: Jiang, Deyang, et al.
Published: (2026)
UItron: Foundational GUI Agent with Advanced Perception and Planning
by: Zeng, Zhixiong, et al.
Published: (2025)
by: Zeng, Zhixiong, et al.
Published: (2025)
UITron-Speech: Towards Automated GUI Agents Based on Speech Instructions
by: Han, Wenkang, et al.
Published: (2025)
by: Han, Wenkang, et al.
Published: (2025)
DocTron-Formula: Generalized Formula Recognition in Complex and Structured Scenarios
by: Zhong, Yufeng, et al.
Published: (2025)
by: Zhong, Yufeng, et al.
Published: (2025)
MobileDreamer: Generative Sketch World Model for GUI Agent
by: Cao, Yilin, et al.
Published: (2026)
by: Cao, Yilin, et al.
Published: (2026)
LightOnOCR: A 1B End-to-End Multilingual Vision-Language Model for State-of-the-Art OCR
by: Taghadouini, Said, et al.
Published: (2026)
by: Taghadouini, Said, et al.
Published: (2026)
End-to-End Vision Tokenizer Tuning
by: Wang, Wenxuan, et al.
Published: (2025)
by: Wang, Wenxuan, et al.
Published: (2025)
GutenOCR: A Grounded Vision-Language Front-End for Documents
by: Heidenreich, Hunter, et al.
Published: (2026)
by: Heidenreich, Hunter, et al.
Published: (2026)
OmniActor: A Generalist GUI and Embodied Agent for 2D&3D Worlds
by: Yang, Longrong, et al.
Published: (2025)
by: Yang, Longrong, et al.
Published: (2025)
General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model
by: Wei, Haoran, et al.
Published: (2024)
by: Wei, Haoran, et al.
Published: (2024)
ORION: A Holistic End-to-End Autonomous Driving Framework by Vision-Language Instructed Action Generation
by: Fu, Haoyu, et al.
Published: (2025)
by: Fu, Haoyu, et al.
Published: (2025)
Qianfan-OCR: A Unified End-to-End Model for Document Intelligence
by: Dong, Daxiang, et al.
Published: (2026)
by: Dong, Daxiang, et al.
Published: (2026)
The End of Manual Decoding: Towards Truly End-to-End Language Models
by: Wang, Zhichao, et al.
Published: (2025)
by: Wang, Zhichao, et al.
Published: (2025)
SOLVE: Synergy of Language-Vision and End-to-End Networks for Autonomous Driving
by: Chen, Xuesong, et al.
Published: (2025)
by: Chen, Xuesong, et al.
Published: (2025)
HERMES: A Holistic End-to-End Risk-Aware Multimodal Embodied System with Vision-Language Models for Long-Tail Autonomous Driving
by: Tang, Weizhe, et al.
Published: (2026)
by: Tang, Weizhe, et al.
Published: (2026)
EVE: Towards End-to-End Video Subtitle Extraction with Vision-Language Models
by: Yu, Haiyang, et al.
Published: (2025)
by: Yu, Haiyang, et al.
Published: (2025)
Senna: Bridging Large Vision-Language Models and End-to-End Autonomous Driving
by: Jiang, Bo, et al.
Published: (2024)
by: Jiang, Bo, et al.
Published: (2024)
DriveMoE: Mixture-of-Experts for Vision-Language-Action Model in End-to-End Autonomous Driving
by: Yang, Zhenjie, et al.
Published: (2025)
by: Yang, Zhenjie, et al.
Published: (2025)
GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
by: Zeng, Aohan, et al.
Published: (2024)
by: Zeng, Aohan, et al.
Published: (2024)
Hint-AD: Holistically Aligned Interpretability in End-to-End Autonomous Driving
by: Ding, Kairui, et al.
Published: (2024)
by: Ding, Kairui, et al.
Published: (2024)
OneVision: An End-to-End Generative Framework for Multi-view E-commerce Vision Search
by: Zheng, Zexin, et al.
Published: (2025)
by: Zheng, Zexin, et al.
Published: (2025)
Towards End-to-End Network Intent Management with Large Language Models
by: Dinh, Lam, et al.
Published: (2025)
by: Dinh, Lam, et al.
Published: (2025)
E2E-MFD: Towards End-to-End Synchronous Multimodal Fusion Detection
by: Zhang, Jiaqing, et al.
Published: (2024)
by: Zhang, Jiaqing, et al.
Published: (2024)
P$^{3}$Nav: End-to-End Perception, Prediction and Planning for Vision-and-Language Navigation
by: Li, Tianfu, et al.
Published: (2026)
by: Li, Tianfu, et al.
Published: (2026)
OpenDriveVLA: Towards End-to-end Autonomous Driving with Large Vision Language Action Model
by: Zhou, Xingcheng, et al.
Published: (2025)
by: Zhou, Xingcheng, et al.
Published: (2025)
VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks
by: Wu, Jiannan, et al.
Published: (2024)
by: Wu, Jiannan, et al.
Published: (2024)
Structured Labeling Enables Faster Vision-Language Models for End-to-End Autonomous Driving
by: Jiang, Hao, et al.
Published: (2025)
by: Jiang, Hao, et al.
Published: (2025)
Towards End-to-End Explainable Facial Action Unit Recognition via Vision-Language Joint Learning
by: Ge, Xuri, et al.
Published: (2024)
by: Ge, Xuri, et al.
Published: (2024)
WiseAD: Knowledge Augmented End-to-End Autonomous Driving with Vision-Language Model
by: Zhang, Songyan, et al.
Published: (2024)
by: Zhang, Songyan, et al.
Published: (2024)
VLM-3D:End-to-End Vision-Language Models for Open-World 3D Perception
by: Chang, Fuhao, et al.
Published: (2025)
by: Chang, Fuhao, et al.
Published: (2025)
Towards End-to-End Quantum Estimation of Non-Hermitian Pseudospectra
by: Yang, Gengzhi, et al.
Published: (2026)
by: Yang, Gengzhi, et al.
Published: (2026)
Towards Fully Decoupled End-to-End Person Search
by: Zhang, Pengcheng, et al.
Published: (2023)
by: Zhang, Pengcheng, et al.
Published: (2023)
Towards End-to-End GPS Localization with Neural Pseudorange Correction
by: Weng, Xu, et al.
Published: (2024)
by: Weng, Xu, et al.
Published: (2024)
Ocean-OCR: Towards General OCR Application via a Vision-Language Model
by: Chen, Song, et al.
Published: (2025)
by: Chen, Song, et al.
Published: (2025)
Agentic Reward Modeling: Verifying GUI Agent via Online Proactive Interaction
by: Cui, Chaoqun, et al.
Published: (2026)
by: Cui, Chaoqun, et al.
Published: (2026)
Similar Items
-
Reading or Reasoning? Format Decoupled Reinforcement Learning for Document OCR
by: Zhong, Yufeng, et al.
Published: (2025) -
VinciCoder: Unifying Multimodal Code Generation via Coarse-to-fine Visual Reinforcement Learning
by: Zhao, Xuanle, et al.
Published: (2025) -
Breaking the SFT Plateau: Multimodal Structured Reinforcement Learning for Chart-to-Code Generation
by: Chen, Lei, et al.
Published: (2025) -
Chart-R1: Chain-of-Thought Supervision and Reinforcement for Advanced Chart Reasoner
by: Chen, Lei, et al.
Published: (2025) -
ScaleTrack: Scaling and back-tracking Automated GUI Agents
by: Huang, Jing, et al.
Published: (2025)