Reading the unreadable: Creating a dataset of 19th century English newspapers using image-to-text language models
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Bourne, Jonathan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CLOCR-C: Context Leveraging OCR Correction with Pre-trained Language Models
von: Bourne, Jonathan
Veröffentlicht: (2024)
von: Bourne, Jonathan
Veröffentlicht: (2024)
Designing large language model prompts to extract scores from messy text: A shared dataset and challenge
von: Thelwall, Mike
Veröffentlicht: (2026)
von: Thelwall, Mike
Veröffentlicht: (2026)
De-identification of clinical free text using natural language processing: A systematic review of current approaches
von: Kovačević, Aleksandar, et al.
Veröffentlicht: (2023)
von: Kovačević, Aleksandar, et al.
Veröffentlicht: (2023)
WithdrarXiv: A Large-Scale Dataset for Retraction Study
von: Rao, Delip, et al.
Veröffentlicht: (2024)
von: Rao, Delip, et al.
Veröffentlicht: (2024)
Can social media provide early warning of retraction? Evidence from critical tweets identified by human annotation and large language models
von: Zheng, Er-Te, et al.
Veröffentlicht: (2024)
von: Zheng, Er-Te, et al.
Veröffentlicht: (2024)
FMMD: A multimodal open peer review dataset based on F1000Research
von: Zhuang, Zhenzhen, et al.
Veröffentlicht: (2026)
von: Zhuang, Zhenzhen, et al.
Veröffentlicht: (2026)
SyROCCo: Enhancing Systematic Reviews using Machine Learning
von: Fang, Zheng, et al.
Veröffentlicht: (2024)
von: Fang, Zheng, et al.
Veröffentlicht: (2024)
Automating Violence Detection and Categorization from Ancient Texts
von: Abdelhalim, Alhassan, et al.
Veröffentlicht: (2025)
von: Abdelhalim, Alhassan, et al.
Veröffentlicht: (2025)
On the Effectiveness of Large Language Models in Automating Categorization of Scientific Texts
von: Shahi, Gautam Kishore, et al.
Veröffentlicht: (2025)
von: Shahi, Gautam Kishore, et al.
Veröffentlicht: (2025)
Falcon 7b for Software Mention Detection in Scholarly Documents
von: Khan, AmeerAli, et al.
Veröffentlicht: (2024)
von: Khan, AmeerAli, et al.
Veröffentlicht: (2024)
Fine-tuning and Prompt Engineering with Cognitive Knowledge Graphs for Scholarly Knowledge Organization
von: Rabby, Gollam, et al.
Veröffentlicht: (2024)
von: Rabby, Gollam, et al.
Veröffentlicht: (2024)
Learning representations of learning representations
von: González-Márquez, Rita, et al.
Veröffentlicht: (2024)
von: González-Márquez, Rita, et al.
Veröffentlicht: (2024)
VTechAGP: An Academic-to-General-Audience Text Paraphrase Dataset and Benchmark Models
von: Cheng, Ming, et al.
Veröffentlicht: (2024)
von: Cheng, Ming, et al.
Veröffentlicht: (2024)
OAG-Bench: A Human-Curated Benchmark for Academic Graph Mining
von: Zhang, Fanjin, et al.
Veröffentlicht: (2024)
von: Zhang, Fanjin, et al.
Veröffentlicht: (2024)
A History of Philosophy in Colombia through Topic Modelling
von: Loaiza, Juan R., et al.
Veröffentlicht: (2024)
von: Loaiza, Juan R., et al.
Veröffentlicht: (2024)
Post-OCR Text Correction for Bulgarian Historical Documents
von: Beshirov, Angel, et al.
Veröffentlicht: (2024)
von: Beshirov, Angel, et al.
Veröffentlicht: (2024)
Machine Learning Research Has Outpaced Its Communication Norms and NeurIPS Should Act
von: Rangarajan, Ajay Mandyam, et al.
Veröffentlicht: (2026)
von: Rangarajan, Ajay Mandyam, et al.
Veröffentlicht: (2026)
Hierarchical Tree-structured Knowledge Graph For Academic Insight Survey
von: Li, Jinghong, et al.
Veröffentlicht: (2024)
von: Li, Jinghong, et al.
Veröffentlicht: (2024)
A Semi-Automatic Approach to Create Large Gender- and Age-Balanced Speaker Corpora: Usefulness of Speaker Diarization & Identification
von: Uro, Rémi, et al.
Veröffentlicht: (2024)
von: Uro, Rémi, et al.
Veröffentlicht: (2024)
Developing ChemDFM as a large language foundation model for chemistry
von: Zhao, Zihan, et al.
Veröffentlicht: (2024)
von: Zhao, Zihan, et al.
Veröffentlicht: (2024)
Interesting Scientific Idea Generation using Knowledge Graphs and LLMs: Evaluations with 100 Research Group Leaders
von: Gu, Xuemei, et al.
Veröffentlicht: (2024)
von: Gu, Xuemei, et al.
Veröffentlicht: (2024)
Combining topic modelling and citation network analysis to study case law from the European Court on Human Rights on the right to respect for private and family life
von: Mohammadi, M., et al.
Veröffentlicht: (2024)
von: Mohammadi, M., et al.
Veröffentlicht: (2024)
Unsupervised extraction of local and global keywords from a single text
von: Aleksanyan, Lida, et al.
Veröffentlicht: (2023)
von: Aleksanyan, Lida, et al.
Veröffentlicht: (2023)
SemEval-2025 Task 5: LLMs4Subjects -- LLM-based Automated Subject Tagging for a National Technical Library's Open-Access Catalog
von: D'Souza, Jennifer, et al.
Veröffentlicht: (2025)
von: D'Souza, Jennifer, et al.
Veröffentlicht: (2025)
Copycats: the many lives of a publicly available medical imaging dataset
von: Jiménez-Sánchez, Amelia, et al.
Veröffentlicht: (2024)
von: Jiménez-Sánchez, Amelia, et al.
Veröffentlicht: (2024)
Storage places in diplomatic texts (7th-13th centuries). Lexical, semantic, and digital investigation
von: Perreaux, Nicolas
Veröffentlicht: (2025)
von: Perreaux, Nicolas
Veröffentlicht: (2025)
SciHorizon: Benchmarking AI-for-Science Readiness from Scientific Data to Large Language Models
von: Qin, Chuan, et al.
Veröffentlicht: (2025)
von: Qin, Chuan, et al.
Veröffentlicht: (2025)
FLAWS: A Benchmark for Error Identification and Localization in Scientific Papers
von: Xi, Sarina, et al.
Veröffentlicht: (2025)
von: Xi, Sarina, et al.
Veröffentlicht: (2025)
TF1-EN-3M: Three Million Synthetic Moral Fables for Training Small, Open Language Models
von: Nadas, Mihai, et al.
Veröffentlicht: (2025)
von: Nadas, Mihai, et al.
Veröffentlicht: (2025)
Publication Trend Analysis and Synthesis via Large Language Model: A Case Study of Engineering in PNAS
von: Smetana, Mason, et al.
Veröffentlicht: (2025)
von: Smetana, Mason, et al.
Veröffentlicht: (2025)
Comparison of Feature Learning Methods for Metadata Extraction from PDF Scholarly Documents
von: Boukhers, Zeyd, et al.
Veröffentlicht: (2025)
von: Boukhers, Zeyd, et al.
Veröffentlicht: (2025)
C$^2$-Cite: Contextual-Aware Citation Generation for Attributed Large Language Models
von: Yu, Yue, et al.
Veröffentlicht: (2025)
von: Yu, Yue, et al.
Veröffentlicht: (2025)
Is ChatGPT Transforming Academics' Writing Style?
von: Geng, Mingmeng, et al.
Veröffentlicht: (2024)
von: Geng, Mingmeng, et al.
Veröffentlicht: (2024)
'Quis custodiet ipsos custodes?' Who will watch the watchmen? On Detecting AI-generated peer-reviews
von: Kumar, Sandeep, et al.
Veröffentlicht: (2024)
von: Kumar, Sandeep, et al.
Veröffentlicht: (2024)
Efficient Systematic Reviews: Literature Filtering with Transformers & Transfer Learning
von: Hawkins, John, et al.
Veröffentlicht: (2024)
von: Hawkins, John, et al.
Veröffentlicht: (2024)
LitLLMs, LLMs for Literature Review: Are we there yet?
von: Agarwal, Shubham, et al.
Veröffentlicht: (2024)
von: Agarwal, Shubham, et al.
Veröffentlicht: (2024)
Chain-of-Factors Paper-Reviewer Matching
von: Zhang, Yu, et al.
Veröffentlicht: (2023)
von: Zhang, Yu, et al.
Veröffentlicht: (2023)
Institutional Books 1.0: A 242B token dataset from Harvard Library's collections, refined for accuracy and usability
von: Cargnelutti, Matteo, et al.
Veröffentlicht: (2025)
von: Cargnelutti, Matteo, et al.
Veröffentlicht: (2025)
Towards understanding evolution of science through language model series
von: Dong, Junjie, et al.
Veröffentlicht: (2024)
von: Dong, Junjie, et al.
Veröffentlicht: (2024)
Meursault as a Data Point
von: Pratap, Abhinav
Veröffentlicht: (2025)
von: Pratap, Abhinav
Veröffentlicht: (2025)
Ähnliche Einträge
-
CLOCR-C: Context Leveraging OCR Correction with Pre-trained Language Models
von: Bourne, Jonathan
Veröffentlicht: (2024) -
Designing large language model prompts to extract scores from messy text: A shared dataset and challenge
von: Thelwall, Mike
Veröffentlicht: (2026) -
De-identification of clinical free text using natural language processing: A systematic review of current approaches
von: Kovačević, Aleksandar, et al.
Veröffentlicht: (2023) -
WithdrarXiv: A Large-Scale Dataset for Retraction Study
von: Rao, Delip, et al.
Veröffentlicht: (2024) -
Can social media provide early warning of retraction? Evidence from critical tweets identified by human annotation and large language models
von: Zheng, Er-Te, et al.
Veröffentlicht: (2024)