TexOCR: Advancing Document OCR Models for Compilable Page-to-LaTeX Reconstruction
Fuente:
arXiv
Guardado en:
| Autores principales: | Wang, Chengye, Fu, Lin, Kuang, Zexi, Zhao, Yilun |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
LaTeX Compilation: Challenges in the Era of LLMs
por: Liu, Tianyou, et al.
Publicado: (2026)
por: Liu, Tianyou, et al.
Publicado: (2026)
SciVer: Evaluating Foundation Models for Multimodal Scientific Claim Verification
por: Wang, Chengye, et al.
Publicado: (2025)
por: Wang, Chengye, et al.
Publicado: (2025)
olmOCR 2: Unit Test Rewards for Document OCR
por: Poznanski, Jake, et al.
Publicado: (2025)
por: Poznanski, Jake, et al.
Publicado: (2025)
Greek2MathTex: A Greek Speech-to-Text Framework for LaTeX Equations Generation
por: Gkritzali, Evangelia, et al.
Publicado: (2024)
por: Gkritzali, Evangelia, et al.
Publicado: (2024)
PreP-OCR: A Complete Pipeline for Document Image Restoration and Enhanced OCR Accuracy
por: Guan, Shuhao, et al.
Publicado: (2025)
por: Guan, Shuhao, et al.
Publicado: (2025)
TeXBLEU: Automatic Metric for Evaluate LaTeX Format
por: Jung, Kyudan, et al.
Publicado: (2024)
por: Jung, Kyudan, et al.
Publicado: (2024)
LaTeXTrans: Structured LaTeX Translation with Multi-Agent Coordination
por: Zhu, Ziming, et al.
Publicado: (2025)
por: Zhu, Ziming, et al.
Publicado: (2025)
Image-to-LaTeX Converter for Mathematical Formulas and Text
por: Gurgurov, Daniil, et al.
Publicado: (2024)
por: Gurgurov, Daniil, et al.
Publicado: (2024)
Multimodal LLMs for OCR, OCR Post-Correction, and Named Entity Recognition in Historical Documents
por: Greif, Gavin, et al.
Publicado: (2025)
por: Greif, Gavin, et al.
Publicado: (2025)
SciMDR: Advancing Scientific Multimodal Document Reasoning
por: Chen, Ziyu, et al.
Publicado: (2026)
por: Chen, Ziyu, et al.
Publicado: (2026)
GLM-OCR Technical Report
por: Duan, Shuaiqi, et al.
Publicado: (2026)
por: Duan, Shuaiqi, et al.
Publicado: (2026)
Typhoon OCR: Open Vision-Language Model For Thai Document Extraction
por: Nonesung, Surapon, et al.
Publicado: (2026)
por: Nonesung, Surapon, et al.
Publicado: (2026)
Reading or Reasoning? Format Decoupled Reinforcement Learning for Document OCR
por: Zhong, Yufeng, et al.
Publicado: (2025)
por: Zhong, Yufeng, et al.
Publicado: (2025)
Can LLMs Generate High-Quality Test Cases for Algorithm Problems? TestCase-Eval: A Systematic Evaluation of Fault Coverage and Exposure
por: Yang, Zheyuan, et al.
Publicado: (2025)
por: Yang, Zheyuan, et al.
Publicado: (2025)
Confidence-Aware Document OCR Error Detection
por: Hemmer, Arthur, et al.
Publicado: (2024)
por: Hemmer, Arthur, et al.
Publicado: (2024)
GlotOCR Bench: OCR Models Still Struggle Beyond a Handful of Unicode Scripts
por: Kargaran, Amir Hossein, et al.
Publicado: (2026)
por: Kargaran, Amir Hossein, et al.
Publicado: (2026)
Seeing Straight: Document Orientation Detection for Efficient OCR
por: Goswami, Suranjan, et al.
Publicado: (2025)
por: Goswami, Suranjan, et al.
Publicado: (2025)
Benchmarking Vision-Language Models on Chinese Ancient Documents: From OCR to Knowledge Reasoning
por: Yu, Haiyang, et al.
Publicado: (2025)
por: Yu, Haiyang, et al.
Publicado: (2025)
Baseer: A Vision-Language Model for Arabic Document-to-Markdown OCR
por: Hennara, Khalil, et al.
Publicado: (2025)
por: Hennara, Khalil, et al.
Publicado: (2025)
KazakhOCR: A Synthetic Benchmark for Evaluating Multimodal Models in Low-Resource Kazakh Script OCR
por: Gagnier, Henry, et al.
Publicado: (2026)
por: Gagnier, Henry, et al.
Publicado: (2026)
PubMed-OCR: PMC Open Access OCR Annotations
por: Heidenreich, Hunter, et al.
Publicado: (2026)
por: Heidenreich, Hunter, et al.
Publicado: (2026)
Advancing Post-OCR Correction: A Comparative Study of Synthetic Data
por: Guan, Shuhao, et al.
Publicado: (2024)
por: Guan, Shuhao, et al.
Publicado: (2024)
OCR Error Post-Correction with LLMs in Historical Documents: No Free Lunches
por: Kanerva, Jenna, et al.
Publicado: (2025)
por: Kanerva, Jenna, et al.
Publicado: (2025)
TeXpert: A Multi-Level Benchmark for Evaluating LaTeX Code Generation by LLMs
por: Kale, Sahil, et al.
Publicado: (2025)
por: Kale, Sahil, et al.
Publicado: (2025)
AI-Friendly LaTeX: Using LaTeX Code as a Knowledge Source for Retrieval-Augmented Generation
por: Verhoeff, Tom
Publicado: (2026)
por: Verhoeff, Tom
Publicado: (2026)
Post-OCR Text Correction for Bulgarian Historical Documents
por: Beshirov, Angel, et al.
Publicado: (2024)
por: Beshirov, Angel, et al.
Publicado: (2024)
$A^2R^2$: Advancing Img2LaTeX Conversion via Visual Reasoning with Attention-Guided Refinement
por: Li, Zhecheng, et al.
Publicado: (2025)
por: Li, Zhecheng, et al.
Publicado: (2025)
Jochre 3 and the Yiddish OCR corpus
por: Urieli, Assaf, et al.
Publicado: (2025)
por: Urieli, Assaf, et al.
Publicado: (2025)
DharmaOCR: Specialized Small Language Models for Structured OCR that outperform Open-Source and Commercial Baselines
por: Cardoso, Gabriel Pimenta de Freitas, et al.
Publicado: (2026)
por: Cardoso, Gabriel Pimenta de Freitas, et al.
Publicado: (2026)
Can Multimodal Foundation Models Understand Schematic Diagrams? An Empirical Study on Information-Seeking QA over Scientific Papers
por: Zhao, Yilun, et al.
Publicado: (2025)
por: Zhao, Yilun, et al.
Publicado: (2025)
olmOCR: Unlocking Trillions of Tokens in PDFs with Vision Language Models
por: Poznanski, Jake, et al.
Publicado: (2025)
por: Poznanski, Jake, et al.
Publicado: (2025)
Investigating OCR-Sensitive Neurons to Improve Entity Recognition in Historical Documents
por: Boros, Emanuela, et al.
Publicado: (2024)
por: Boros, Emanuela, et al.
Publicado: (2024)
FastOCR: Dynamic Visual Fixation via KV Cache Pruning for Efficient Document Parsing
por: Tang, Zihan, et al.
Publicado: (2026)
por: Tang, Zihan, et al.
Publicado: (2026)
OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models
por: Liu, Yuliang, et al.
Publicado: (2023)
por: Liu, Yuliang, et al.
Publicado: (2023)
CC-OCR V2: Benchmarking Large Multimodal Models for Literacy in Real-world Document Processing
por: Xu, Zhipeng, et al.
Publicado: (2026)
por: Xu, Zhipeng, et al.
Publicado: (2026)
Advances and Limitations in Open Source Arabic-Script OCR: A Case Study
por: Kiessling, Benjamin, et al.
Publicado: (2024)
por: Kiessling, Benjamin, et al.
Publicado: (2024)
Automated LaTeX Code Generation from Handwritten Math Expressions Using Vision Transformer
por: Sundararaj, Jayaprakash, et al.
Publicado: (2024)
por: Sundararaj, Jayaprakash, et al.
Publicado: (2024)
QID: Efficient Query-Informed ViTs in Data-Scarce Regimes for OCR-free Visual Document Understanding
por: Le, Binh M., et al.
Publicado: (2025)
por: Le, Binh M., et al.
Publicado: (2025)
Towards Deployable OCR models for Indic languages
por: Mathew, Minesh, et al.
Publicado: (2022)
por: Mathew, Minesh, et al.
Publicado: (2022)
Improving OCR for Historical Texts of Multiple Languages
por: Westerdijk, Hylke, et al.
Publicado: (2025)
por: Westerdijk, Hylke, et al.
Publicado: (2025)
Ejemplares similares
-
LaTeX Compilation: Challenges in the Era of LLMs
por: Liu, Tianyou, et al.
Publicado: (2026) -
SciVer: Evaluating Foundation Models for Multimodal Scientific Claim Verification
por: Wang, Chengye, et al.
Publicado: (2025) -
olmOCR 2: Unit Test Rewards for Document OCR
por: Poznanski, Jake, et al.
Publicado: (2025) -
Greek2MathTex: A Greek Speech-to-Text Framework for LaTeX Equations Generation
por: Gkritzali, Evangelia, et al.
Publicado: (2024) -
PreP-OCR: A Complete Pipeline for Document Image Restoration and Enhanced OCR Accuracy
por: Guan, Shuhao, et al.
Publicado: (2025)