Post-OCR Text Correction for Bulgarian Historical Documents
Fuente:
arXiv
Saved in:
| Main Authors: | Beshirov, Angel, Dobreva, Milena, Dimitrov, Dimitar, Hardalov, Momchil, Koychev, Ivan, Nakov, Preslav |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Detecting Check-Worthy Claims in Political Debates, Speeches, and Interviews Using Audio Data
by: Ivanov, Petar, et al.
Published: (2023)
by: Ivanov, Petar, et al.
Published: (2023)
Multimodal LLMs for OCR, OCR Post-Correction, and Named Entity Recognition in Historical Documents
by: Greif, Gavin, et al.
Published: (2025)
by: Greif, Gavin, et al.
Published: (2025)
PubMed-OCR: PMC Open Access OCR Annotations
by: Heidenreich, Hunter, et al.
Published: (2026)
by: Heidenreich, Hunter, et al.
Published: (2026)
FMI_SU_Yotkova_Kastreva at SemEval-2026 Task 13: Lightweight Detection of LLM-Generated Code via Stylometric Signals
by: Yotkova, Elitsa, et al.
Published: (2026)
by: Yotkova, Elitsa, et al.
Published: (2026)
Automating Violence Detection and Categorization from Ancient Texts
by: Abdelhalim, Alhassan, et al.
Published: (2025)
by: Abdelhalim, Alhassan, et al.
Published: (2025)
On the Effectiveness of Large Language Models in Automating Categorization of Scientific Texts
by: Shahi, Gautam Kishore, et al.
Published: (2025)
by: Shahi, Gautam Kishore, et al.
Published: (2025)
Unlocking the Archives: Using Large Language Models to Transcribe Handwritten Historical Documents
by: Humphries, Mark, et al.
Published: (2024)
by: Humphries, Mark, et al.
Published: (2024)
Falcon 7b for Software Mention Detection in Scholarly Documents
by: Khan, AmeerAli, et al.
Published: (2024)
by: Khan, AmeerAli, et al.
Published: (2024)
VTechAGP: An Academic-to-General-Audience Text Paraphrase Dataset and Benchmark Models
by: Cheng, Ming, et al.
Published: (2024)
by: Cheng, Ming, et al.
Published: (2024)
CLOCR-C: Context Leveraging OCR Correction with Pre-trained Language Models
by: Bourne, Jonathan
Published: (2024)
by: Bourne, Jonathan
Published: (2024)
Historical Ink: 19th Century Latin American Spanish Newspaper Corpus with LLM OCR Correction
by: Manrique-Gómez, Laura, et al.
Published: (2024)
by: Manrique-Gómez, Laura, et al.
Published: (2024)
Fine-tuning and Prompt Engineering with Cognitive Knowledge Graphs for Scholarly Knowledge Organization
by: Rabby, Gollam, et al.
Published: (2024)
by: Rabby, Gollam, et al.
Published: (2024)
WithdrarXiv: A Large-Scale Dataset for Retraction Study
by: Rao, Delip, et al.
Published: (2024)
by: Rao, Delip, et al.
Published: (2024)
Learning representations of learning representations
by: González-Márquez, Rita, et al.
Published: (2024)
by: González-Márquez, Rita, et al.
Published: (2024)
Reading the unreadable: Creating a dataset of 19th century English newspapers using image-to-text language models
by: Bourne, Jonathan
Published: (2025)
by: Bourne, Jonathan
Published: (2025)
OAG-Bench: A Human-Curated Benchmark for Academic Graph Mining
by: Zhang, Fanjin, et al.
Published: (2024)
by: Zhang, Fanjin, et al.
Published: (2024)
A History of Philosophy in Colombia through Topic Modelling
by: Loaiza, Juan R., et al.
Published: (2024)
by: Loaiza, Juan R., et al.
Published: (2024)
Machine Learning Research Has Outpaced Its Communication Norms and NeurIPS Should Act
by: Rangarajan, Ajay Mandyam, et al.
Published: (2026)
by: Rangarajan, Ajay Mandyam, et al.
Published: (2026)
Hierarchical Tree-structured Knowledge Graph For Academic Insight Survey
by: Li, Jinghong, et al.
Published: (2024)
by: Li, Jinghong, et al.
Published: (2024)
Comparison of Feature Learning Methods for Metadata Extraction from PDF Scholarly Documents
by: Boukhers, Zeyd, et al.
Published: (2025)
by: Boukhers, Zeyd, et al.
Published: (2025)
SyROCCo: Enhancing Systematic Reviews using Machine Learning
by: Fang, Zheng, et al.
Published: (2024)
by: Fang, Zheng, et al.
Published: (2024)
Named Entity Recognition of Historical Texts via Large Language Model
by: Zhang, Shibingfeng, et al.
Published: (2025)
by: Zhang, Shibingfeng, et al.
Published: (2025)
A Survey of OCR Evaluation Methods and Metrics and the Invisibility of Historical Documents
by: Beyene, Fitsum Sileshi, et al.
Published: (2026)
by: Beyene, Fitsum Sileshi, et al.
Published: (2026)
Combining topic modelling and citation network analysis to study case law from the European Court on Human Rights on the right to respect for private and family life
by: Mohammadi, M., et al.
Published: (2024)
by: Mohammadi, M., et al.
Published: (2024)
Is ChatGPT Transforming Academics' Writing Style?
by: Geng, Mingmeng, et al.
Published: (2024)
by: Geng, Mingmeng, et al.
Published: (2024)
SciHorizon: Benchmarking AI-for-Science Readiness from Scientific Data to Large Language Models
by: Qin, Chuan, et al.
Published: (2025)
by: Qin, Chuan, et al.
Published: (2025)
FLAWS: A Benchmark for Error Identification and Localization in Scientific Papers
by: Xi, Sarina, et al.
Published: (2025)
by: Xi, Sarina, et al.
Published: (2025)
'Quis custodiet ipsos custodes?' Who will watch the watchmen? On Detecting AI-generated peer-reviews
by: Kumar, Sandeep, et al.
Published: (2024)
by: Kumar, Sandeep, et al.
Published: (2024)
Efficient Systematic Reviews: Literature Filtering with Transformers & Transfer Learning
by: Hawkins, John, et al.
Published: (2024)
by: Hawkins, John, et al.
Published: (2024)
LitLLMs, LLMs for Literature Review: Are we there yet?
by: Agarwal, Shubham, et al.
Published: (2024)
by: Agarwal, Shubham, et al.
Published: (2024)
Chain-of-Factors Paper-Reviewer Matching
by: Zhang, Yu, et al.
Published: (2023)
by: Zhang, Yu, et al.
Published: (2023)
TF1-EN-3M: Three Million Synthetic Moral Fables for Training Small, Open Language Models
by: Nadas, Mihai, et al.
Published: (2025)
by: Nadas, Mihai, et al.
Published: (2025)
SemEval-2025 Task 5: LLMs4Subjects -- LLM-based Automated Subject Tagging for a National Technical Library's Open-Access Catalog
by: D'Souza, Jennifer, et al.
Published: (2025)
by: D'Souza, Jennifer, et al.
Published: (2025)
Publication Trend Analysis and Synthesis via Large Language Model: A Case Study of Engineering in PNAS
by: Smetana, Mason, et al.
Published: (2025)
by: Smetana, Mason, et al.
Published: (2025)
Can social media provide early warning of retraction? Evidence from critical tweets identified by human annotation and large language models
by: Zheng, Er-Te, et al.
Published: (2024)
by: Zheng, Er-Te, et al.
Published: (2024)
FMMD: A multimodal open peer review dataset based on F1000Research
by: Zhuang, Zhenzhen, et al.
Published: (2026)
by: Zhuang, Zhenzhen, et al.
Published: (2026)
C$^2$-Cite: Contextual-Aware Citation Generation for Attributed Large Language Models
by: Yu, Yue, et al.
Published: (2025)
by: Yu, Yue, et al.
Published: (2025)
Interesting Scientific Idea Generation using Knowledge Graphs and LLMs: Evaluations with 100 Research Group Leaders
by: Gu, Xuemei, et al.
Published: (2024)
by: Gu, Xuemei, et al.
Published: (2024)
Labeling Case Similarity based on Co-Citation of Legal Articles in Judgment Documents with Empirical Dispute-Based Evaluation
by: Liu, Chao-Lin, et al.
Published: (2025)
by: Liu, Chao-Lin, et al.
Published: (2025)
EXAMS-V: A Multi-Discipline Multilingual Multimodal Exam Benchmark for Evaluating Vision Language Models
by: Das, Rocktim Jyoti, et al.
Published: (2024)
by: Das, Rocktim Jyoti, et al.
Published: (2024)
Similar Items
-
Detecting Check-Worthy Claims in Political Debates, Speeches, and Interviews Using Audio Data
by: Ivanov, Petar, et al.
Published: (2023) -
Multimodal LLMs for OCR, OCR Post-Correction, and Named Entity Recognition in Historical Documents
by: Greif, Gavin, et al.
Published: (2025) -
PubMed-OCR: PMC Open Access OCR Annotations
by: Heidenreich, Hunter, et al.
Published: (2026) -
FMI_SU_Yotkova_Kastreva at SemEval-2026 Task 13: Lightweight Detection of LLM-Generated Code via Stylometric Signals
by: Yotkova, Elitsa, et al.
Published: (2026) -
Automating Violence Detection and Categorization from Ancient Texts
by: Abdelhalim, Alhassan, et al.
Published: (2025)