PEaCE: A Chemistry-Oriented Dataset for Optical Character Recognition on Scientific Documents
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Nan, Heaton, Connor, Okonsky, Sean Timothy, Mitra, Prasenjit, Toraman, Hilal Ezgi |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Effect of pyrolysis operating conditions on the catalytic co‐pyrolysis of low‐density polyethylene and polyethylene terephthalate with zeolite catalysts
by: Sean Timothy Okonsky, et al.
Published: (2024)
by: Sean Timothy Okonsky, et al.
Published: (2024)
LlamaTurk: Adapting Open-Source Generative Large Language Models for Low-Resource Language
by: Toraman, Cagri
Published: (2024)
by: Toraman, Cagri
Published: (2024)
E-ARMOR: Edge case Assessment and Review of Multilingual Optical Character Recognition
by: Gupta, Aryan, et al.
Published: (2025)
by: Gupta, Aryan, et al.
Published: (2025)
LOCR: Location-Guided Transformer for Optical Character Recognition
by: Sun, Yu, et al.
Published: (2024)
by: Sun, Yu, et al.
Published: (2024)
Rethinking Genomic Modeling Through Optical Character Recognition
by: Xiang, Hongxin, et al.
Published: (2026)
by: Xiang, Hongxin, et al.
Published: (2026)
Qalam : A Multimodal LLM for Arabic Optical Character and Handwriting Recognition
by: Bhatia, Gagan, et al.
Published: (2024)
by: Bhatia, Gagan, et al.
Published: (2024)
Towards Efficient Methods in Medical Question Answering using Knowledge Graph Embeddings
by: Sengupta, Saptarshi, et al.
Published: (2024)
by: Sengupta, Saptarshi, et al.
Published: (2024)
FIBER: A Multilingual Evaluation Resource for Factual Inference Bias
by: Munis, Evren Ayberk, et al.
Published: (2025)
by: Munis, Evren Ayberk, et al.
Published: (2025)
SiReRAG: Indexing Similar and Related Information for Multihop Reasoning
by: Zhang, Nan, et al.
Published: (2024)
by: Zhang, Nan, et al.
Published: (2024)
BIRDTurk: Adaptation of the BIRD Text-to-SQL Dataset to Turkish
by: Aktaş, Burak, et al.
Published: (2026)
by: Aktaş, Burak, et al.
Published: (2026)
Evaluating the Quality of Benchmark Datasets for Low-Resource Languages: A Case Study on Turkish
by: Cengiz, Ayşe Aysu, et al.
Published: (2025)
by: Cengiz, Ayşe Aysu, et al.
Published: (2025)
Pragya: An AI-Based Semantic Recommendation System for Sanskrit Subhasitas
by: Raorane, Tanisha, et al.
Published: (2026)
by: Raorane, Tanisha, et al.
Published: (2026)
IRPAPERS: A Visual Document Benchmark for Scientific Retrieval and Question Answering
by: Shorten, Connor, et al.
Published: (2026)
by: Shorten, Connor, et al.
Published: (2026)
From Language Models over Tokens to Language Models over Characters
by: Vieira, Tim, et al.
Published: (2024)
by: Vieira, Tim, et al.
Published: (2024)
Mining Large Language Models for Low-Resource Language Data: Comparing Elicitation Strategies for Hausa and Fongbe
by: Adjovi, Mahounan Pericles, et al.
Published: (2026)
by: Adjovi, Mahounan Pericles, et al.
Published: (2026)
When Does Data Augmentation Help? Evaluating LLM and Back-Translation Methods for Hausa and Fongbe NLP
by: Adjovi, Mahounan Pericles, et al.
Published: (2026)
by: Adjovi, Mahounan Pericles, et al.
Published: (2026)
T1: A Tool-Oriented Conversational Dataset for Multi-Turn Agentic Planning
by: Chakraborty, Amartya, et al.
Published: (2025)
by: Chakraborty, Amartya, et al.
Published: (2025)
MASSW: A New Dataset and Benchmark Tasks for AI-Assisted Scientific Workflows
by: Zhang, Xingjian, et al.
Published: (2024)
by: Zhang, Xingjian, et al.
Published: (2024)
DAGverse: Building Document-Grounded Semantic DAGs from Scientific Papers
by: Wan, Shu, et al.
Published: (2026)
by: Wan, Shu, et al.
Published: (2026)
Iterative Auto-Annotation for Scientific Named Entity Recognition Using BERT-Based Models
by: Gupta, Kartik
Published: (2025)
by: Gupta, Kartik
Published: (2025)
BookWorm: A Dataset for Character Description and Analysis
by: Papoudakis, Argyrios, et al.
Published: (2024)
by: Papoudakis, Argyrios, et al.
Published: (2024)
MOOSE-Chem: Large Language Models for Rediscovering Unseen Chemistry Scientific Hypotheses
by: Yang, Zonglin, et al.
Published: (2024)
by: Yang, Zonglin, et al.
Published: (2024)
Pearce's Characterisation in an Epistemic Domain
by: Su, Ezgi Iraz
Published: (2025)
by: Su, Ezgi Iraz
Published: (2025)
AI Managed Emergency Documentation with a Pretrained Model
by: Menzies, David, et al.
Published: (2024)
by: Menzies, David, et al.
Published: (2024)
Advancing Scientific Text Classification: Fine-Tuned Models with Dataset Expansion and Hard-Voting
by: Rostam, Zhyar Rzgar K, et al.
Published: (2025)
by: Rostam, Zhyar Rzgar K, et al.
Published: (2025)
Query-driven Document-level Scientific Evidence Extraction from Biomedical Studies
by: Pronesti, Massimiliano, et al.
Published: (2025)
by: Pronesti, Massimiliano, et al.
Published: (2025)
Killkan: The Automatic Speech Recognition Dataset for Kichwa with Morphosyntactic Information
by: Taguchi, Chihiro, et al.
Published: (2024)
by: Taguchi, Chihiro, et al.
Published: (2024)
Making Task-Oriented Dialogue Datasets More Natural by Synthetically Generating Indirect User Requests
by: Mannekote, Amogh, et al.
Published: (2024)
by: Mannekote, Amogh, et al.
Published: (2024)
JMultiWOZ: A Large-Scale Japanese Multi-Domain Task-Oriented Dialogue Dataset
by: Ohashi, Atsumoto, et al.
Published: (2024)
by: Ohashi, Atsumoto, et al.
Published: (2024)
Leveraging Large Language Models for Rare Disease Named Entity Recognition
by: Xi, Nan Miles, et al.
Published: (2025)
by: Xi, Nan Miles, et al.
Published: (2025)
Reddit-Impacts: A Named Entity Recognition Dataset for Analyzing Clinical and Social Effects of Substance Use Derived from Social Media
by: Ge, Yao, et al.
Published: (2024)
by: Ge, Yao, et al.
Published: (2024)
MDCR: A Dataset for Multi-Document Conditional Reasoning
by: Chen, Peter Baile, et al.
Published: (2024)
by: Chen, Peter Baile, et al.
Published: (2024)
Document-as-Image Representations Fall Short for Scientific Retrieval
by: Khalighinejad, Ghazal, et al.
Published: (2026)
by: Khalighinejad, Ghazal, et al.
Published: (2026)
Uncertainty-Aware Fusion: An Ensemble Framework for Mitigating Hallucinations in Large Language Models
by: Dey, Prasenjit, et al.
Published: (2025)
by: Dey, Prasenjit, et al.
Published: (2025)
Generating Visual Stories with Grounded and Coreferent Characters
by: Liu, Danyang, et al.
Published: (2024)
by: Liu, Danyang, et al.
Published: (2024)
Can Large Language Models Generate Effective Datasets for Emotion Recognition in Conversations?
by: Kaplan, Burak Can, et al.
Published: (2025)
by: Kaplan, Burak Can, et al.
Published: (2025)
OCRTurk: A Comprehensive OCR Benchmark for Turkish
by: Yılmaz, Deniz, et al.
Published: (2026)
by: Yılmaz, Deniz, et al.
Published: (2026)
MME-RAG: Multi-Manager-Expert Retrieval-Augmented Generation for Fine-Grained Entity Recognition in Task-Oriented Dialogues
by: Xue, Liang, et al.
Published: (2025)
by: Xue, Liang, et al.
Published: (2025)
Investigating OCR-Sensitive Neurons to Improve Entity Recognition in Historical Documents
by: Boros, Emanuela, et al.
Published: (2024)
by: Boros, Emanuela, et al.
Published: (2024)
Triples and Knowledge-Infused Embeddings for Clustering and Classification of Scientific Documents
by: Arcan, Mihael
Published: (2025)
by: Arcan, Mihael
Published: (2025)
Similar Items
-
Effect of pyrolysis operating conditions on the catalytic co‐pyrolysis of low‐density polyethylene and polyethylene terephthalate with zeolite catalysts
by: Sean Timothy Okonsky, et al.
Published: (2024) -
LlamaTurk: Adapting Open-Source Generative Large Language Models for Low-Resource Language
by: Toraman, Cagri
Published: (2024) -
E-ARMOR: Edge case Assessment and Review of Multilingual Optical Character Recognition
by: Gupta, Aryan, et al.
Published: (2025) -
LOCR: Location-Guided Transformer for Optical Character Recognition
by: Sun, Yu, et al.
Published: (2024) -
Rethinking Genomic Modeling Through Optical Character Recognition
by: Xiang, Hongxin, et al.
Published: (2026)