PEaCE: A Chemistry-Oriented Dataset for Optical Character Recognition on Scientific Documents
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhang, Nan, Heaton, Connor, Okonsky, Sean Timothy, Mitra, Prasenjit, Toraman, Hilal Ezgi |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Effect of pyrolysis operating conditions on the catalytic co‐pyrolysis of low‐density polyethylene and polyethylene terephthalate with zeolite catalysts
por: Sean Timothy Okonsky, et al.
Publicado: (2024)
por: Sean Timothy Okonsky, et al.
Publicado: (2024)
LlamaTurk: Adapting Open-Source Generative Large Language Models for Low-Resource Language
por: Toraman, Cagri
Publicado: (2024)
por: Toraman, Cagri
Publicado: (2024)
E-ARMOR: Edge case Assessment and Review of Multilingual Optical Character Recognition
por: Gupta, Aryan, et al.
Publicado: (2025)
por: Gupta, Aryan, et al.
Publicado: (2025)
LOCR: Location-Guided Transformer for Optical Character Recognition
por: Sun, Yu, et al.
Publicado: (2024)
por: Sun, Yu, et al.
Publicado: (2024)
Rethinking Genomic Modeling Through Optical Character Recognition
por: Xiang, Hongxin, et al.
Publicado: (2026)
por: Xiang, Hongxin, et al.
Publicado: (2026)
Qalam : A Multimodal LLM for Arabic Optical Character and Handwriting Recognition
por: Bhatia, Gagan, et al.
Publicado: (2024)
por: Bhatia, Gagan, et al.
Publicado: (2024)
Towards Efficient Methods in Medical Question Answering using Knowledge Graph Embeddings
por: Sengupta, Saptarshi, et al.
Publicado: (2024)
por: Sengupta, Saptarshi, et al.
Publicado: (2024)
FIBER: A Multilingual Evaluation Resource for Factual Inference Bias
por: Munis, Evren Ayberk, et al.
Publicado: (2025)
por: Munis, Evren Ayberk, et al.
Publicado: (2025)
SiReRAG: Indexing Similar and Related Information for Multihop Reasoning
por: Zhang, Nan, et al.
Publicado: (2024)
por: Zhang, Nan, et al.
Publicado: (2024)
BIRDTurk: Adaptation of the BIRD Text-to-SQL Dataset to Turkish
por: Aktaş, Burak, et al.
Publicado: (2026)
por: Aktaş, Burak, et al.
Publicado: (2026)
Evaluating the Quality of Benchmark Datasets for Low-Resource Languages: A Case Study on Turkish
por: Cengiz, Ayşe Aysu, et al.
Publicado: (2025)
por: Cengiz, Ayşe Aysu, et al.
Publicado: (2025)
Pragya: An AI-Based Semantic Recommendation System for Sanskrit Subhasitas
por: Raorane, Tanisha, et al.
Publicado: (2026)
por: Raorane, Tanisha, et al.
Publicado: (2026)
IRPAPERS: A Visual Document Benchmark for Scientific Retrieval and Question Answering
por: Shorten, Connor, et al.
Publicado: (2026)
por: Shorten, Connor, et al.
Publicado: (2026)
From Language Models over Tokens to Language Models over Characters
por: Vieira, Tim, et al.
Publicado: (2024)
por: Vieira, Tim, et al.
Publicado: (2024)
T1: A Tool-Oriented Conversational Dataset for Multi-Turn Agentic Planning
por: Chakraborty, Amartya, et al.
Publicado: (2025)
por: Chakraborty, Amartya, et al.
Publicado: (2025)
Mining Large Language Models for Low-Resource Language Data: Comparing Elicitation Strategies for Hausa and Fongbe
por: Adjovi, Mahounan Pericles, et al.
Publicado: (2026)
por: Adjovi, Mahounan Pericles, et al.
Publicado: (2026)
When Does Data Augmentation Help? Evaluating LLM and Back-Translation Methods for Hausa and Fongbe NLP
por: Adjovi, Mahounan Pericles, et al.
Publicado: (2026)
por: Adjovi, Mahounan Pericles, et al.
Publicado: (2026)
MASSW: A New Dataset and Benchmark Tasks for AI-Assisted Scientific Workflows
por: Zhang, Xingjian, et al.
Publicado: (2024)
por: Zhang, Xingjian, et al.
Publicado: (2024)
DAGverse: Building Document-Grounded Semantic DAGs from Scientific Papers
por: Wan, Shu, et al.
Publicado: (2026)
por: Wan, Shu, et al.
Publicado: (2026)
Iterative Auto-Annotation for Scientific Named Entity Recognition Using BERT-Based Models
por: Gupta, Kartik
Publicado: (2025)
por: Gupta, Kartik
Publicado: (2025)
BookWorm: A Dataset for Character Description and Analysis
por: Papoudakis, Argyrios, et al.
Publicado: (2024)
por: Papoudakis, Argyrios, et al.
Publicado: (2024)
MOOSE-Chem: Large Language Models for Rediscovering Unseen Chemistry Scientific Hypotheses
por: Yang, Zonglin, et al.
Publicado: (2024)
por: Yang, Zonglin, et al.
Publicado: (2024)
Pearce's Characterisation in an Epistemic Domain
por: Su, Ezgi Iraz
Publicado: (2025)
por: Su, Ezgi Iraz
Publicado: (2025)
AI Managed Emergency Documentation with a Pretrained Model
por: Menzies, David, et al.
Publicado: (2024)
por: Menzies, David, et al.
Publicado: (2024)
Advancing Scientific Text Classification: Fine-Tuned Models with Dataset Expansion and Hard-Voting
por: Rostam, Zhyar Rzgar K, et al.
Publicado: (2025)
por: Rostam, Zhyar Rzgar K, et al.
Publicado: (2025)
Query-driven Document-level Scientific Evidence Extraction from Biomedical Studies
por: Pronesti, Massimiliano, et al.
Publicado: (2025)
por: Pronesti, Massimiliano, et al.
Publicado: (2025)
Killkan: The Automatic Speech Recognition Dataset for Kichwa with Morphosyntactic Information
por: Taguchi, Chihiro, et al.
Publicado: (2024)
por: Taguchi, Chihiro, et al.
Publicado: (2024)
Making Task-Oriented Dialogue Datasets More Natural by Synthetically Generating Indirect User Requests
por: Mannekote, Amogh, et al.
Publicado: (2024)
por: Mannekote, Amogh, et al.
Publicado: (2024)
JMultiWOZ: A Large-Scale Japanese Multi-Domain Task-Oriented Dialogue Dataset
por: Ohashi, Atsumoto, et al.
Publicado: (2024)
por: Ohashi, Atsumoto, et al.
Publicado: (2024)
Leveraging Large Language Models for Rare Disease Named Entity Recognition
por: Xi, Nan Miles, et al.
Publicado: (2025)
por: Xi, Nan Miles, et al.
Publicado: (2025)
Reddit-Impacts: A Named Entity Recognition Dataset for Analyzing Clinical and Social Effects of Substance Use Derived from Social Media
por: Ge, Yao, et al.
Publicado: (2024)
por: Ge, Yao, et al.
Publicado: (2024)
MDCR: A Dataset for Multi-Document Conditional Reasoning
por: Chen, Peter Baile, et al.
Publicado: (2024)
por: Chen, Peter Baile, et al.
Publicado: (2024)
Document-as-Image Representations Fall Short for Scientific Retrieval
por: Khalighinejad, Ghazal, et al.
Publicado: (2026)
por: Khalighinejad, Ghazal, et al.
Publicado: (2026)
Uncertainty-Aware Fusion: An Ensemble Framework for Mitigating Hallucinations in Large Language Models
por: Dey, Prasenjit, et al.
Publicado: (2025)
por: Dey, Prasenjit, et al.
Publicado: (2025)
Generating Visual Stories with Grounded and Coreferent Characters
por: Liu, Danyang, et al.
Publicado: (2024)
por: Liu, Danyang, et al.
Publicado: (2024)
Can Large Language Models Generate Effective Datasets for Emotion Recognition in Conversations?
por: Kaplan, Burak Can, et al.
Publicado: (2025)
por: Kaplan, Burak Can, et al.
Publicado: (2025)
MME-RAG: Multi-Manager-Expert Retrieval-Augmented Generation for Fine-Grained Entity Recognition in Task-Oriented Dialogues
por: Xue, Liang, et al.
Publicado: (2025)
por: Xue, Liang, et al.
Publicado: (2025)
OCRTurk: A Comprehensive OCR Benchmark for Turkish
por: Yılmaz, Deniz, et al.
Publicado: (2026)
por: Yılmaz, Deniz, et al.
Publicado: (2026)
Investigating OCR-Sensitive Neurons to Improve Entity Recognition in Historical Documents
por: Boros, Emanuela, et al.
Publicado: (2024)
por: Boros, Emanuela, et al.
Publicado: (2024)
Triples and Knowledge-Infused Embeddings for Clustering and Classification of Scientific Documents
por: Arcan, Mihael
Publicado: (2025)
por: Arcan, Mihael
Publicado: (2025)
Ejemplares similares
-
Effect of pyrolysis operating conditions on the catalytic co‐pyrolysis of low‐density polyethylene and polyethylene terephthalate with zeolite catalysts
por: Sean Timothy Okonsky, et al.
Publicado: (2024) -
LlamaTurk: Adapting Open-Source Generative Large Language Models for Low-Resource Language
por: Toraman, Cagri
Publicado: (2024) -
E-ARMOR: Edge case Assessment and Review of Multilingual Optical Character Recognition
por: Gupta, Aryan, et al.
Publicado: (2025) -
LOCR: Location-Guided Transformer for Optical Character Recognition
por: Sun, Yu, et al.
Publicado: (2024) -
Rethinking Genomic Modeling Through Optical Character Recognition
por: Xiang, Hongxin, et al.
Publicado: (2026)