Reading or Reasoning? Format Decoupled Reinforcement Learning for Document OCR
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhong, Yufeng, Chen, Lei, Zeng, Zhixiong, Zhao, Xuanle, Jiang, Deyang, Zheng, Liming, Huang, Jing, Qiu, Haibo, Shi, Peng, Yang, Siqi, Ma, Lin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
VinciCoder: Unifying Multimodal Code Generation via Coarse-to-fine Visual Reinforcement Learning
von: Zhao, Xuanle, et al.
Veröffentlicht: (2025)
von: Zhao, Xuanle, et al.
Veröffentlicht: (2025)
OCRVerse: Towards Holistic OCR in End-to-End Vision-Language Models
von: Zhong, Yufeng, et al.
Veröffentlicht: (2026)
von: Zhong, Yufeng, et al.
Veröffentlicht: (2026)
TreeCUA: Efficiently Scaling GUI Automation with Tree-Structured Verifiable Evolution
von: Jiang, Deyang, et al.
Veröffentlicht: (2026)
von: Jiang, Deyang, et al.
Veröffentlicht: (2026)
Breaking the SFT Plateau: Multimodal Structured Reinforcement Learning for Chart-to-Code Generation
von: Chen, Lei, et al.
Veröffentlicht: (2025)
von: Chen, Lei, et al.
Veröffentlicht: (2025)
Chart-R1: Chain-of-Thought Supervision and Reinforcement for Advanced Chart Reasoner
von: Chen, Lei, et al.
Veröffentlicht: (2025)
von: Chen, Lei, et al.
Veröffentlicht: (2025)
Metis-SPECS: Decoupling Multimodal Learning via Self-distilled Preference-based Cold Start
von: Chen, Kun, et al.
Veröffentlicht: (2025)
von: Chen, Kun, et al.
Veröffentlicht: (2025)
Perceptual-Evidence Anchored Reinforced Learning for Multimodal Reasoning
von: Zhang, Chi, et al.
Veröffentlicht: (2025)
von: Zhang, Chi, et al.
Veröffentlicht: (2025)
DocTron-Formula: Generalized Formula Recognition in Complex and Structured Scenarios
von: Zhong, Yufeng, et al.
Veröffentlicht: (2025)
von: Zhong, Yufeng, et al.
Veröffentlicht: (2025)
OmniActor: A Generalist GUI and Embodied Agent for 2D&3D Worlds
von: Yang, Longrong, et al.
Veröffentlicht: (2025)
von: Yang, Longrong, et al.
Veröffentlicht: (2025)
Learning When to Look: A Disentangled Curriculum for Strategic Perception in Multimodal Reasoning
von: Yang, Siqi, et al.
Veröffentlicht: (2025)
von: Yang, Siqi, et al.
Veröffentlicht: (2025)
Length-Unbiased Sequence Policy Optimization: Revealing and Controlling Response Length Variation in RLVR
von: Liu, Fanfan, et al.
Veröffentlicht: (2026)
von: Liu, Fanfan, et al.
Veröffentlicht: (2026)
DeepSketcher: Internalizing Visual Manipulation for Multimodal Reasoning
von: Zhang, Chi, et al.
Veröffentlicht: (2025)
von: Zhang, Chi, et al.
Veröffentlicht: (2025)
Flexible Entropy Control in RLVR with a Gradient-Preserving Perspective
von: Chen, Kun, et al.
Veröffentlicht: (2026)
von: Chen, Kun, et al.
Veröffentlicht: (2026)
STAGE: Stable and Generalizable GRPO for Autoregressive Image Generation
von: Ma, Xiaoxiao, et al.
Veröffentlicht: (2025)
von: Ma, Xiaoxiao, et al.
Veröffentlicht: (2025)
UItron: Foundational GUI Agent with Advanced Perception and Planning
von: Zeng, Zhixiong, et al.
Veröffentlicht: (2025)
von: Zeng, Zhixiong, et al.
Veröffentlicht: (2025)
UITron-Speech: Towards Automated GUI Agents Based on Speech Instructions
von: Han, Wenkang, et al.
Veröffentlicht: (2025)
von: Han, Wenkang, et al.
Veröffentlicht: (2025)
SimpleOCR: Rendering Visualized Questions to Teach MLLMs to Read
von: Peng, Yibo, et al.
Veröffentlicht: (2026)
von: Peng, Yibo, et al.
Veröffentlicht: (2026)
MobileDreamer: Generative Sketch World Model for GUI Agent
von: Cao, Yilin, et al.
Veröffentlicht: (2026)
von: Cao, Yilin, et al.
Veröffentlicht: (2026)
Metis-HOME: Hybrid Optimized Mixture-of-Experts for Multimodal Reasoning
von: Lan, Xiaohan, et al.
Veröffentlicht: (2025)
von: Lan, Xiaohan, et al.
Veröffentlicht: (2025)
GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling
von: Li, Siqi, et al.
Veröffentlicht: (2025)
von: Li, Siqi, et al.
Veröffentlicht: (2025)
Benchmarking Vision-Language Models on Chinese Ancient Documents: From OCR to Knowledge Reasoning
von: Yu, Haiyang, et al.
Veröffentlicht: (2025)
von: Yu, Haiyang, et al.
Veröffentlicht: (2025)
olmOCR 2: Unit Test Rewards for Document OCR
von: Poznanski, Jake, et al.
Veröffentlicht: (2025)
von: Poznanski, Jake, et al.
Veröffentlicht: (2025)
LogicOCR: Do Your Large Multimodal Models Excel at Logical Reasoning on Text-Rich Images?
von: Ye, Maoyuan, et al.
Veröffentlicht: (2025)
von: Ye, Maoyuan, et al.
Veröffentlicht: (2025)
Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues?
von: He, Haibin, et al.
Veröffentlicht: (2025)
von: He, Haibin, et al.
Veröffentlicht: (2025)
Multimodal OCR: Parse Anything from Documents
von: Zheng, Handong, et al.
Veröffentlicht: (2026)
von: Zheng, Handong, et al.
Veröffentlicht: (2026)
Towards Better & Faster Autoregressive Image Generation: From the Perspective of Entropy
von: Ma, Xiaoxiao, et al.
Veröffentlicht: (2025)
von: Ma, Xiaoxiao, et al.
Veröffentlicht: (2025)
Read Quietly, Think Aloud: Decoupling Comprehension and Reasoning in LLMs
von: Wang, Yuanxin, et al.
Veröffentlicht: (2025)
von: Wang, Yuanxin, et al.
Veröffentlicht: (2025)
PreP-OCR: A Complete Pipeline for Document Image Restoration and Enhanced OCR Accuracy
von: Guan, Shuhao, et al.
Veröffentlicht: (2025)
von: Guan, Shuhao, et al.
Veröffentlicht: (2025)
ChemVLR: Prioritizing Reasoning in Perception for Chemical Vision-Language Understanding
von: Zhao, Xuanle, et al.
Veröffentlicht: (2026)
von: Zhao, Xuanle, et al.
Veröffentlicht: (2026)
What to Format and How: A Benchmark and Workflow Approach for Document Formatting
von: Rao, Shihao, et al.
Veröffentlicht: (2026)
von: Rao, Shihao, et al.
Veröffentlicht: (2026)
Confidence-Aware Document OCR Error Detection
von: Hemmer, Arthur, et al.
Veröffentlicht: (2024)
von: Hemmer, Arthur, et al.
Veröffentlicht: (2024)
Decoupling Task-Solving and Output Formatting in LLM Generation
von: Deng, Haikang, et al.
Veröffentlicht: (2025)
von: Deng, Haikang, et al.
Veröffentlicht: (2025)
Hybrid Latent Reasoning with Decoupled Policy Optimization
von: Cheng, Tao, et al.
Veröffentlicht: (2026)
von: Cheng, Tao, et al.
Veröffentlicht: (2026)
Seeing Straight: Document Orientation Detection for Efficient OCR
von: Goswami, Suranjan, et al.
Veröffentlicht: (2025)
von: Goswami, Suranjan, et al.
Veröffentlicht: (2025)
Metis-RISE: RL Incentivizes and SFT Enhances Multimodal Reasoning Model Learning
von: Qiu, Haibo, et al.
Veröffentlicht: (2025)
von: Qiu, Haibo, et al.
Veröffentlicht: (2025)
TexOCR: Advancing Document OCR Models for Compilable Page-to-LaTeX Reconstruction
von: Wang, Chengye, et al.
Veröffentlicht: (2026)
von: Wang, Chengye, et al.
Veröffentlicht: (2026)
Combining OCR Models for Reading Early Modern Printed Books
von: Seuret, Mathias, et al.
Veröffentlicht: (2023)
von: Seuret, Mathias, et al.
Veröffentlicht: (2023)
GITA: Graph to Visual and Textual Integration for Vision-Language Graph Reasoning
von: Wei, Yanbin, et al.
Veröffentlicht: (2024)
von: Wei, Yanbin, et al.
Veröffentlicht: (2024)
Mitigating Training Imbalance in LLM Fine-Tuning via Selective Parameter Merging
von: Ju, Yiming, et al.
Veröffentlicht: (2024)
von: Ju, Yiming, et al.
Veröffentlicht: (2024)
Detached Skip-Links and $R$-Probe: Decoupling Feature Aggregation from Gradient Propagation for MLLM OCR
von: Yuan, Ziye, et al.
Veröffentlicht: (2026)
von: Yuan, Ziye, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
VinciCoder: Unifying Multimodal Code Generation via Coarse-to-fine Visual Reinforcement Learning
von: Zhao, Xuanle, et al.
Veröffentlicht: (2025) -
OCRVerse: Towards Holistic OCR in End-to-End Vision-Language Models
von: Zhong, Yufeng, et al.
Veröffentlicht: (2026) -
TreeCUA: Efficiently Scaling GUI Automation with Tree-Structured Verifiable Evolution
von: Jiang, Deyang, et al.
Veröffentlicht: (2026) -
Breaking the SFT Plateau: Multimodal Structured Reinforcement Learning for Chart-to-Code Generation
von: Chen, Lei, et al.
Veröffentlicht: (2025) -
Chart-R1: Chain-of-Thought Supervision and Reinforcement for Advanced Chart Reasoner
von: Chen, Lei, et al.
Veröffentlicht: (2025)