TexOCR: Advancing Document OCR Models for Compilable Page-to-LaTeX Reconstruction
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Chengye, Fu, Lin, Kuang, Zexi, Zhao, Yilun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LaTeX Compilation: Challenges in the Era of LLMs
von: Liu, Tianyou, et al.
Veröffentlicht: (2026)
von: Liu, Tianyou, et al.
Veröffentlicht: (2026)
SciVer: Evaluating Foundation Models for Multimodal Scientific Claim Verification
von: Wang, Chengye, et al.
Veröffentlicht: (2025)
von: Wang, Chengye, et al.
Veröffentlicht: (2025)
olmOCR 2: Unit Test Rewards for Document OCR
von: Poznanski, Jake, et al.
Veröffentlicht: (2025)
von: Poznanski, Jake, et al.
Veröffentlicht: (2025)
Greek2MathTex: A Greek Speech-to-Text Framework for LaTeX Equations Generation
von: Gkritzali, Evangelia, et al.
Veröffentlicht: (2024)
von: Gkritzali, Evangelia, et al.
Veröffentlicht: (2024)
PreP-OCR: A Complete Pipeline for Document Image Restoration and Enhanced OCR Accuracy
von: Guan, Shuhao, et al.
Veröffentlicht: (2025)
von: Guan, Shuhao, et al.
Veröffentlicht: (2025)
TeXBLEU: Automatic Metric for Evaluate LaTeX Format
von: Jung, Kyudan, et al.
Veröffentlicht: (2024)
von: Jung, Kyudan, et al.
Veröffentlicht: (2024)
LaTeXTrans: Structured LaTeX Translation with Multi-Agent Coordination
von: Zhu, Ziming, et al.
Veröffentlicht: (2025)
von: Zhu, Ziming, et al.
Veröffentlicht: (2025)
Image-to-LaTeX Converter for Mathematical Formulas and Text
von: Gurgurov, Daniil, et al.
Veröffentlicht: (2024)
von: Gurgurov, Daniil, et al.
Veröffentlicht: (2024)
Multimodal LLMs for OCR, OCR Post-Correction, and Named Entity Recognition in Historical Documents
von: Greif, Gavin, et al.
Veröffentlicht: (2025)
von: Greif, Gavin, et al.
Veröffentlicht: (2025)
SciMDR: Advancing Scientific Multimodal Document Reasoning
von: Chen, Ziyu, et al.
Veröffentlicht: (2026)
von: Chen, Ziyu, et al.
Veröffentlicht: (2026)
GLM-OCR Technical Report
von: Duan, Shuaiqi, et al.
Veröffentlicht: (2026)
von: Duan, Shuaiqi, et al.
Veröffentlicht: (2026)
Typhoon OCR: Open Vision-Language Model For Thai Document Extraction
von: Nonesung, Surapon, et al.
Veröffentlicht: (2026)
von: Nonesung, Surapon, et al.
Veröffentlicht: (2026)
Reading or Reasoning? Format Decoupled Reinforcement Learning for Document OCR
von: Zhong, Yufeng, et al.
Veröffentlicht: (2025)
von: Zhong, Yufeng, et al.
Veröffentlicht: (2025)
Can LLMs Generate High-Quality Test Cases for Algorithm Problems? TestCase-Eval: A Systematic Evaluation of Fault Coverage and Exposure
von: Yang, Zheyuan, et al.
Veröffentlicht: (2025)
von: Yang, Zheyuan, et al.
Veröffentlicht: (2025)
Confidence-Aware Document OCR Error Detection
von: Hemmer, Arthur, et al.
Veröffentlicht: (2024)
von: Hemmer, Arthur, et al.
Veröffentlicht: (2024)
GlotOCR Bench: OCR Models Still Struggle Beyond a Handful of Unicode Scripts
von: Kargaran, Amir Hossein, et al.
Veröffentlicht: (2026)
von: Kargaran, Amir Hossein, et al.
Veröffentlicht: (2026)
Seeing Straight: Document Orientation Detection for Efficient OCR
von: Goswami, Suranjan, et al.
Veröffentlicht: (2025)
von: Goswami, Suranjan, et al.
Veröffentlicht: (2025)
Benchmarking Vision-Language Models on Chinese Ancient Documents: From OCR to Knowledge Reasoning
von: Yu, Haiyang, et al.
Veröffentlicht: (2025)
von: Yu, Haiyang, et al.
Veröffentlicht: (2025)
Baseer: A Vision-Language Model for Arabic Document-to-Markdown OCR
von: Hennara, Khalil, et al.
Veröffentlicht: (2025)
von: Hennara, Khalil, et al.
Veröffentlicht: (2025)
KazakhOCR: A Synthetic Benchmark for Evaluating Multimodal Models in Low-Resource Kazakh Script OCR
von: Gagnier, Henry, et al.
Veröffentlicht: (2026)
von: Gagnier, Henry, et al.
Veröffentlicht: (2026)
PubMed-OCR: PMC Open Access OCR Annotations
von: Heidenreich, Hunter, et al.
Veröffentlicht: (2026)
von: Heidenreich, Hunter, et al.
Veröffentlicht: (2026)
Advancing Post-OCR Correction: A Comparative Study of Synthetic Data
von: Guan, Shuhao, et al.
Veröffentlicht: (2024)
von: Guan, Shuhao, et al.
Veröffentlicht: (2024)
OCR Error Post-Correction with LLMs in Historical Documents: No Free Lunches
von: Kanerva, Jenna, et al.
Veröffentlicht: (2025)
von: Kanerva, Jenna, et al.
Veröffentlicht: (2025)
TeXpert: A Multi-Level Benchmark for Evaluating LaTeX Code Generation by LLMs
von: Kale, Sahil, et al.
Veröffentlicht: (2025)
von: Kale, Sahil, et al.
Veröffentlicht: (2025)
AI-Friendly LaTeX: Using LaTeX Code as a Knowledge Source for Retrieval-Augmented Generation
von: Verhoeff, Tom
Veröffentlicht: (2026)
von: Verhoeff, Tom
Veröffentlicht: (2026)
Post-OCR Text Correction for Bulgarian Historical Documents
von: Beshirov, Angel, et al.
Veröffentlicht: (2024)
von: Beshirov, Angel, et al.
Veröffentlicht: (2024)
$A^2R^2$: Advancing Img2LaTeX Conversion via Visual Reasoning with Attention-Guided Refinement
von: Li, Zhecheng, et al.
Veröffentlicht: (2025)
von: Li, Zhecheng, et al.
Veröffentlicht: (2025)
Jochre 3 and the Yiddish OCR corpus
von: Urieli, Assaf, et al.
Veröffentlicht: (2025)
von: Urieli, Assaf, et al.
Veröffentlicht: (2025)
DharmaOCR: Specialized Small Language Models for Structured OCR that outperform Open-Source and Commercial Baselines
von: Cardoso, Gabriel Pimenta de Freitas, et al.
Veröffentlicht: (2026)
von: Cardoso, Gabriel Pimenta de Freitas, et al.
Veröffentlicht: (2026)
Can Multimodal Foundation Models Understand Schematic Diagrams? An Empirical Study on Information-Seeking QA over Scientific Papers
von: Zhao, Yilun, et al.
Veröffentlicht: (2025)
von: Zhao, Yilun, et al.
Veröffentlicht: (2025)
olmOCR: Unlocking Trillions of Tokens in PDFs with Vision Language Models
von: Poznanski, Jake, et al.
Veröffentlicht: (2025)
von: Poznanski, Jake, et al.
Veröffentlicht: (2025)
Investigating OCR-Sensitive Neurons to Improve Entity Recognition in Historical Documents
von: Boros, Emanuela, et al.
Veröffentlicht: (2024)
von: Boros, Emanuela, et al.
Veröffentlicht: (2024)
FastOCR: Dynamic Visual Fixation via KV Cache Pruning for Efficient Document Parsing
von: Tang, Zihan, et al.
Veröffentlicht: (2026)
von: Tang, Zihan, et al.
Veröffentlicht: (2026)
OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models
von: Liu, Yuliang, et al.
Veröffentlicht: (2023)
von: Liu, Yuliang, et al.
Veröffentlicht: (2023)
CC-OCR V2: Benchmarking Large Multimodal Models for Literacy in Real-world Document Processing
von: Xu, Zhipeng, et al.
Veröffentlicht: (2026)
von: Xu, Zhipeng, et al.
Veröffentlicht: (2026)
Advances and Limitations in Open Source Arabic-Script OCR: A Case Study
von: Kiessling, Benjamin, et al.
Veröffentlicht: (2024)
von: Kiessling, Benjamin, et al.
Veröffentlicht: (2024)
Automated LaTeX Code Generation from Handwritten Math Expressions Using Vision Transformer
von: Sundararaj, Jayaprakash, et al.
Veröffentlicht: (2024)
von: Sundararaj, Jayaprakash, et al.
Veröffentlicht: (2024)
QID: Efficient Query-Informed ViTs in Data-Scarce Regimes for OCR-free Visual Document Understanding
von: Le, Binh M., et al.
Veröffentlicht: (2025)
von: Le, Binh M., et al.
Veröffentlicht: (2025)
Towards Deployable OCR models for Indic languages
von: Mathew, Minesh, et al.
Veröffentlicht: (2022)
von: Mathew, Minesh, et al.
Veröffentlicht: (2022)
Improving OCR for Historical Texts of Multiple Languages
von: Westerdijk, Hylke, et al.
Veröffentlicht: (2025)
von: Westerdijk, Hylke, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
LaTeX Compilation: Challenges in the Era of LLMs
von: Liu, Tianyou, et al.
Veröffentlicht: (2026) -
SciVer: Evaluating Foundation Models for Multimodal Scientific Claim Verification
von: Wang, Chengye, et al.
Veröffentlicht: (2025) -
olmOCR 2: Unit Test Rewards for Document OCR
von: Poznanski, Jake, et al.
Veröffentlicht: (2025) -
Greek2MathTex: A Greek Speech-to-Text Framework for LaTeX Equations Generation
von: Gkritzali, Evangelia, et al.
Veröffentlicht: (2024) -
PreP-OCR: A Complete Pipeline for Document Image Restoration and Enhanced OCR Accuracy
von: Guan, Shuhao, et al.
Veröffentlicht: (2025)