The Patrologia Graeca Corpus: OCR, Annotation, and Open Release of Noisy Nineteenth-Century Polytonic Greek Editions
Fuente:
arXiv
Saved in:
| Main Authors: | Vidal-Gorène, Chahan, Kindt, Bastien |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Under-resourced studies of under-resourced languages: lemmatization and POS-tagging with LLM annotators for historical Armenian, Georgian, Greek and Syriac
by: Vidal-Gorène, Chahan, et al.
Published: (2026)
by: Vidal-Gorène, Chahan, et al.
Published: (2026)
Timing In stand-up Comedy: Text, Audio, Laughter, Kinesics (TIC-TALK): Pipeline and Database for the Multimodal Study of Comedic Timing
by: Zribi, Yaelle, et al.
Published: (2026)
by: Zribi, Yaelle, et al.
Published: (2026)
Logios : An open source Greek Polytonic Optical Character Recognition system
by: Konstantinos, Perifanos, et al.
Published: (2025)
by: Konstantinos, Perifanos, et al.
Published: (2025)
PubMed-OCR: PMC Open Access OCR Annotations
by: Heidenreich, Hunter, et al.
Published: (2026)
by: Heidenreich, Hunter, et al.
Published: (2026)
Reading or Guessing? Visual Grounding Failures of Vision-Language Models for OCR in Ancient Greek Editions
by: Karamolegkou, Antonia, et al.
Published: (2026)
by: Karamolegkou, Antonia, et al.
Published: (2026)
OCR-Quality: A Human-Annotated Dataset for OCR Quality Assessment
by: Zhang, Yulong
Published: (2025)
by: Zhang, Yulong
Published: (2025)
Callico: a Versatile Open-Source Document Image Annotation Platform
by: Kermorvant, Christopher, et al.
Published: (2024)
by: Kermorvant, Christopher, et al.
Published: (2024)
Noisy Annotations in Semantic Segmentation
by: Kimhi, Moshe, et al.
Published: (2024)
by: Kimhi, Moshe, et al.
Published: (2024)
DocParseNet: Advanced Semantic Segmentation and OCR Embeddings for Efficient Scanned Document Annotation
by: Mohammadshirazi, Ahmad, et al.
Published: (2024)
by: Mohammadshirazi, Ahmad, et al.
Published: (2024)
Active Learning with a Noisy Annotator
by: Shafir, Netta, et al.
Published: (2025)
by: Shafir, Netta, et al.
Published: (2025)
Holistic Reliability Propagation: Decoupling Annotation and Prediction for Robust Noisy-Label
by: Mao, Jingyang, et al.
Published: (2026)
by: Mao, Jingyang, et al.
Published: (2026)
AtlasOCR: Building the First Open-Source Darija OCR Model with Vision Language Models
by: Momayiz, Imane, et al.
Published: (2026)
by: Momayiz, Imane, et al.
Published: (2026)
OCR-Agent: Agentic OCR with Capability and Memory Reflection
by: Wen, Shimin, et al.
Published: (2026)
by: Wen, Shimin, et al.
Published: (2026)
OmniOCR: Generalist OCR for Ethnic Minority Languages
by: Liu, Bonan, et al.
Published: (2026)
by: Liu, Bonan, et al.
Published: (2026)
Cross-Lingual SynthDocs: A Large-Scale Synthetic Corpus for Any to Arabic OCR and Document Understanding
by: Al-Homoud, Haneen, et al.
Published: (2025)
by: Al-Homoud, Haneen, et al.
Published: (2025)
dopanim: A Dataset of Doppelganger Animals with Noisy Annotations from Multiple Humans
by: Herde, Marek, et al.
Published: (2024)
by: Herde, Marek, et al.
Published: (2024)
Collaborative Group: Composed Image Retrieval via Consensus Learning from Noisy Annotations
by: Zhang, Xu, et al.
Published: (2023)
by: Zhang, Xu, et al.
Published: (2023)
Label Filling via Mixed Supervision for Medical Image Segmentation from Noisy Annotations
by: Li, Ming, et al.
Published: (2024)
by: Li, Ming, et al.
Published: (2024)
Agentar-Fin-OCR
by: Qian, Siyi, et al.
Published: (2026)
by: Qian, Siyi, et al.
Published: (2026)
When Good OCR Is Not Enough: Benchmarking OCR Robustness for Retrieval-Augmented Generation
by: Sun, Lin, et al.
Published: (2026)
by: Sun, Lin, et al.
Published: (2026)
OCR Hinders RAG: Evaluating the Cascading Impact of OCR on Retrieval-Augmented Generation
by: Zhang, Junyuan, et al.
Published: (2024)
by: Zhang, Junyuan, et al.
Published: (2024)
DharmaOCR: Specialized Small Language Models for Structured OCR that outperform Open-Source and Commercial Baselines
by: Cardoso, Gabriel Pimenta de Freitas, et al.
Published: (2026)
by: Cardoso, Gabriel Pimenta de Freitas, et al.
Published: (2026)
Advances and Limitations in Open Source Arabic-Script OCR: A Case Study
by: Kiessling, Benjamin, et al.
Published: (2024)
by: Kiessling, Benjamin, et al.
Published: (2024)
MME-VideoOCR: Evaluating OCR-Based Capabilities of Multimodal LLMs in Video Scenarios
by: Shi, Yang, et al.
Published: (2025)
by: Shi, Yang, et al.
Published: (2025)
Ocean-OCR: Towards General OCR Application via a Vision-Language Model
by: Chen, Song, et al.
Published: (2025)
by: Chen, Song, et al.
Published: (2025)
TC-OCR: TableCraft OCR for Efficient Detection & Recognition of Table Structure & Content
by: Anand, Avinash, et al.
Published: (2024)
by: Anand, Avinash, et al.
Published: (2024)
olmOCR 2: Unit Test Rewards for Document OCR
by: Poznanski, Jake, et al.
Published: (2025)
by: Poznanski, Jake, et al.
Published: (2025)
The Socface Project: Large-Scale Collection, Processing, and Analysis of a Century of French Censuses
by: Boillet, Mélodie, et al.
Published: (2024)
by: Boillet, Mélodie, et al.
Published: (2024)
How to Efficiently Annotate Images for Best-Performing Deep Learning Based Segmentation Models: An Empirical Study with Weak and Noisy Annotations and Segment Anything Model
by: Zhang, Yixin, et al.
Published: (2023)
by: Zhang, Yixin, et al.
Published: (2023)
ABot-OCR Technical Report
by: Jiang, Kaitao, et al.
Published: (2026)
by: Jiang, Kaitao, et al.
Published: (2026)
CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy
by: Yang, Zhibo, et al.
Published: (2024)
by: Yang, Zhibo, et al.
Published: (2024)
Incorporating Crowdsourced Annotator Distributions into Ensemble Modeling to Improve Classification Trustworthiness for Ancient Greek Papyri
by: West, Graham, et al.
Published: (2022)
by: West, Graham, et al.
Published: (2022)
Error Patterns in Historical OCR: A Comparative Analysis of TrOCR and a Vision-Language Model
by: Vesalainen, Ari, et al.
Published: (2026)
by: Vesalainen, Ari, et al.
Published: (2026)
General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model
by: Wei, Haoran, et al.
Published: (2024)
by: Wei, Haoran, et al.
Published: (2024)
From Veracity to Diffusion: Adressing Operational Challenges in Moving From Fake-News Detection to Information Disorders
by: Savatteri, Francesco Paolo, et al.
Published: (2025)
by: Savatteri, Francesco Paolo, et al.
Published: (2025)
DODO: Discrete OCR Diffusion Models
by: Man, Sean, et al.
Published: (2026)
by: Man, Sean, et al.
Published: (2026)
An Empirical Study of Scaling Law for OCR
by: Rang, Miao, et al.
Published: (2023)
by: Rang, Miao, et al.
Published: (2023)
LightOnOCR: A 1B End-to-End Multilingual Vision-Language Model for State-of-the-Art OCR
by: Taghadouini, Said, et al.
Published: (2026)
by: Taghadouini, Said, et al.
Published: (2026)
Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues?
by: He, Haibin, et al.
Published: (2025)
by: He, Haibin, et al.
Published: (2025)
OpenBox: Annotate Any Bounding Boxes in 3D
by: Lee, In-Jae, et al.
Published: (2025)
by: Lee, In-Jae, et al.
Published: (2025)
Similar Items
-
Under-resourced studies of under-resourced languages: lemmatization and POS-tagging with LLM annotators for historical Armenian, Georgian, Greek and Syriac
by: Vidal-Gorène, Chahan, et al.
Published: (2026) -
Timing In stand-up Comedy: Text, Audio, Laughter, Kinesics (TIC-TALK): Pipeline and Database for the Multimodal Study of Comedic Timing
by: Zribi, Yaelle, et al.
Published: (2026) -
Logios : An open source Greek Polytonic Optical Character Recognition system
by: Konstantinos, Perifanos, et al.
Published: (2025) -
PubMed-OCR: PMC Open Access OCR Annotations
by: Heidenreich, Hunter, et al.
Published: (2026) -
Reading or Guessing? Visual Grounding Failures of Vision-Language Models for OCR in Ancient Greek Editions
by: Karamolegkou, Antonia, et al.
Published: (2026)