Enriching Historical Records: An OCR and AI-Driven Approach for Database Integration
Fuente:
arXiv
Guardado en:
| Autores principales: | Abedi, Zahra, van Dijk, Richard M. K., Wijnholds, Gijs, Verhoef, Tessa |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Dataset Creation for Visual Entailment using Generative AI
por: Reijtenbach, Rob, et al.
Publicado: (2025)
por: Reijtenbach, Rob, et al.
Publicado: (2025)
Investigating OCR-Sensitive Neurons to Improve Entity Recognition in Historical Documents
por: Boros, Emanuela, et al.
Publicado: (2024)
por: Boros, Emanuela, et al.
Publicado: (2024)
Multimodal LLMs for OCR, OCR Post-Correction, and Named Entity Recognition in Historical Documents
por: Greif, Gavin, et al.
Publicado: (2025)
por: Greif, Gavin, et al.
Publicado: (2025)
The Curious Case of Representational Alignment: Unravelling Visio-Linguistic Tasks in Emergent Communication
por: Kouwenhoven, Tom, et al.
Publicado: (2024)
por: Kouwenhoven, Tom, et al.
Publicado: (2024)
Detecting Hallucinations in Large Language Models via Internal Attention Divergence Signals
por: van Dijk, Gijs
Publicado: (2026)
por: van Dijk, Gijs
Publicado: (2026)
A Review of Challenges in Speech-based Conversational AI for Elderly Care
por: Klaassen, Willemijn, et al.
Publicado: (2024)
por: Klaassen, Willemijn, et al.
Publicado: (2024)
Towards Semantically Enriched Embeddings for Knowledge Graph Completion
por: Alam, Mehwish, et al.
Publicado: (2023)
por: Alam, Mehwish, et al.
Publicado: (2023)
Probing Vision-Language Understanding through the Visual Entailment Task: promises and pitfalls
por: Pitta, Elena, et al.
Publicado: (2025)
por: Pitta, Elena, et al.
Publicado: (2025)
Unlocking Electronic Health Records: A Hybrid Graph RAG Approach to Safe Clinical AI for Patient QA
por: Thio, Samuel, et al.
Publicado: (2025)
por: Thio, Samuel, et al.
Publicado: (2025)
Evaluating LLMs for Historical Document OCR: A Methodological Framework for Digital Humanities
por: Levchenko, Maria
Publicado: (2025)
por: Levchenko, Maria
Publicado: (2025)
OCRTurk: A Comprehensive OCR Benchmark for Turkish
por: Yılmaz, Deniz, et al.
Publicado: (2026)
por: Yılmaz, Deniz, et al.
Publicado: (2026)
EnrichEvent: Enriching Social Data with Contextual Information for Emerging Event Extraction
por: Esfahani, Mohammadali Sefidi, et al.
Publicado: (2023)
por: Esfahani, Mohammadali Sefidi, et al.
Publicado: (2023)
Cultural Perspectives and Expectations for Generative AI: A Global Survey Approach
por: van Liemt, Erin, et al.
Publicado: (2026)
por: van Liemt, Erin, et al.
Publicado: (2026)
Measuring Political Preferences in AI Systems: An Integrative Approach
por: Rozado, David
Publicado: (2025)
por: Rozado, David
Publicado: (2025)
LLMs for Translation: Historical, Low-Resourced Languages and Contemporary AI Models
por: Tekgurler, Merve
Publicado: (2025)
por: Tekgurler, Merve
Publicado: (2025)
Spanish TrOCR: Leveraging Transfer Learning for Language Adaptation
por: Lauar, Filipe, et al.
Publicado: (2024)
por: Lauar, Filipe, et al.
Publicado: (2024)
Coordinates from Context: Using LLMs to Ground Complex Location References
por: Masis, Tessa, et al.
Publicado: (2025)
por: Masis, Tessa, et al.
Publicado: (2025)
Where on Earth Do Users Say They Are?: Geo-Entity Linking for Noisy Multilingual User Input
por: Masis, Tessa, et al.
Publicado: (2024)
por: Masis, Tessa, et al.
Publicado: (2024)
Enriched BERT Embeddings for Scholarly Publication Classification
por: Wolff, Benjamin, et al.
Publicado: (2024)
por: Wolff, Benjamin, et al.
Publicado: (2024)
Integrating Large Language Models with Human Expertise for Disease Detection in Electronic Health Records
por: Pan, Jie, et al.
Publicado: (2025)
por: Pan, Jie, et al.
Publicado: (2025)
EnrichIndex: Using LLMs to Enrich Retrieval Indices Offline
por: Chen, Peter Baile, et al.
Publicado: (2025)
por: Chen, Peter Baile, et al.
Publicado: (2025)
Categorical Vector Space Semantics for Lambek Calculus with a Relevant Modality
por: McPheat, Lachlan, et al.
Publicado: (2020)
por: McPheat, Lachlan, et al.
Publicado: (2020)
Enhancing AI-Driven Education: Integrating Cognitive Frameworks, Linguistic Feedback Analysis, and Ethical Considerations for Improved Content Generation
por: Yaacoub, Antoun, et al.
Publicado: (2025)
por: Yaacoub, Antoun, et al.
Publicado: (2025)
EQ-5D Classification Using Biomedical Entity-Enriched Pre-trained Language Models and Multiple Instance Learning
por: Rostam, Zhyar Rzgar K, et al.
Publicado: (2026)
por: Rostam, Zhyar Rzgar K, et al.
Publicado: (2026)
Is Open Source the Future of AI? A Data-Driven Approach
por: Vake, Domen, et al.
Publicado: (2025)
por: Vake, Domen, et al.
Publicado: (2025)
Developing and Evaluating an AI-Assisted Prediction Model for Unplanned Intensive Care Admissions following Elective Neurosurgery using Natural Language Processing within an Electronic Healthcare Record System
por: Ive, Julia, et al.
Publicado: (2025)
por: Ive, Julia, et al.
Publicado: (2025)
Quantifying Label-Induced Bias in Large Language Model Self- and Cross-Evaluations
por: Saraf, Muskan, et al.
Publicado: (2025)
por: Saraf, Muskan, et al.
Publicado: (2025)
DharmaOCR: Specialized Small Language Models for Structured OCR that outperform Open-Source and Commercial Baselines
por: Cardoso, Gabriel Pimenta de Freitas, et al.
Publicado: (2026)
por: Cardoso, Gabriel Pimenta de Freitas, et al.
Publicado: (2026)
System-2 Mathematical Reasoning via Enriched Instruction Tuning
por: Cai, Huanqia, et al.
Publicado: (2024)
por: Cai, Huanqia, et al.
Publicado: (2024)
From Medical Records to Diagnostic Dialogues: A Clinical-Grounded Approach and Dataset for Psychiatric Comorbidity
por: Wan, Tianxi, et al.
Publicado: (2025)
por: Wan, Tianxi, et al.
Publicado: (2025)
OCR or Not? Rethinking Document Information Extraction in the MLLMs Era with Real-World Large-Scale Datasets
por: Shen, Jiyuan, et al.
Publicado: (2026)
por: Shen, Jiyuan, et al.
Publicado: (2026)
A Blueprint for AI-Driven Software Quality: Integrating LLMs with Established Standards
por: Patil, Avinash
Publicado: (2025)
por: Patil, Avinash
Publicado: (2025)
Integrating Personality into Digital Humans: A Review of LLM-Driven Approaches for Virtual Reality
por: Brito, Iago Alves, et al.
Publicado: (2025)
por: Brito, Iago Alves, et al.
Publicado: (2025)
Decompose, Enrich, and Extract! Schema-aware Event Extraction using LLMs
por: Shiri, Fatemeh, et al.
Publicado: (2024)
por: Shiri, Fatemeh, et al.
Publicado: (2024)
Hybrid OCR-LLM Framework for Enterprise-Scale Document Information Extraction Under Copy-heavy Task
por: Wang, Zilong, et al.
Publicado: (2025)
por: Wang, Zilong, et al.
Publicado: (2025)
CorIL: Towards Enriching Indian Language to Indian Language Parallel Corpora and Machine Translation Systems
por: Bhattacharjee, Soham, et al.
Publicado: (2025)
por: Bhattacharjee, Soham, et al.
Publicado: (2025)
Simulating the Emergence of Differential Case Marking with Communicating Neural-Network Agents
por: Lian, Yuchen, et al.
Publicado: (2025)
por: Lian, Yuchen, et al.
Publicado: (2025)
What does Kiki look like? Cross-modal associations between speech sounds and visual shapes in vision-and-language models
por: Verhoef, Tessa, et al.
Publicado: (2024)
por: Verhoef, Tessa, et al.
Publicado: (2024)
Searching for Structure: Investigating Emergent Communication with Large Language Models
por: Kouwenhoven, Tom, et al.
Publicado: (2024)
por: Kouwenhoven, Tom, et al.
Publicado: (2024)
NeLLCom-X: A Comprehensive Neural-Agent Framework to Simulate Language Learning and Group Communication
por: Lian, Yuchen, et al.
Publicado: (2024)
por: Lian, Yuchen, et al.
Publicado: (2024)
Ejemplares similares
-
Dataset Creation for Visual Entailment using Generative AI
por: Reijtenbach, Rob, et al.
Publicado: (2025) -
Investigating OCR-Sensitive Neurons to Improve Entity Recognition in Historical Documents
por: Boros, Emanuela, et al.
Publicado: (2024) -
Multimodal LLMs for OCR, OCR Post-Correction, and Named Entity Recognition in Historical Documents
por: Greif, Gavin, et al.
Publicado: (2025) -
The Curious Case of Representational Alignment: Unravelling Visio-Linguistic Tasks in Emergent Communication
por: Kouwenhoven, Tom, et al.
Publicado: (2024) -
Detecting Hallucinations in Large Language Models via Internal Attention Divergence Signals
por: van Dijk, Gijs
Publicado: (2026)