Benchmarking Agentic Systems in Automated Scientific Information Extraction with ChemX

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Vepreva, Anastasia, Razlivina, Julia, Eremeeva, Maria, Gubina, Nina, Orlova, Anastasia, Dmitrenko, Aleksei, Kapranova, Ksenya, Jyakhwo, Susan, Vasilev, Nikita, Sarkisyan, Arsen, Chernyshov, Ivan Yu., Vinogradov, Vladimir, Dmitrenko, Andrei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911187767132160
author Vepreva, Anastasia
Razlivina, Julia
Eremeeva, Maria
Gubina, Nina
Orlova, Anastasia
Dmitrenko, Aleksei
Kapranova, Ksenya
Jyakhwo, Susan
Vasilev, Nikita
Sarkisyan, Arsen
Chernyshov, Ivan Yu.
Vinogradov, Vladimir
Dmitrenko, Andrei
author_facet Vepreva, Anastasia
Razlivina, Julia
Eremeeva, Maria
Gubina, Nina
Orlova, Anastasia
Dmitrenko, Aleksei
Kapranova, Ksenya
Jyakhwo, Susan
Vasilev, Nikita
Sarkisyan, Arsen
Chernyshov, Ivan Yu.
Vinogradov, Vladimir
Dmitrenko, Andrei
contents The emergence of agent-based systems represents a significant advancement in artificial intelligence, with growing applications in automated data extraction. However, chemical information extraction remains a formidable challenge due to the inherent heterogeneity of chemical data. Current agent-based approaches, both general-purpose and domain-specific, exhibit limited performance in this domain. To address this gap, we present ChemX, a comprehensive collection of 10 manually curated and domain-expert-validated datasets focusing on nanomaterials and small molecules. These datasets are designed to rigorously evaluate and enhance automated extraction methodologies in chemistry. To demonstrate their utility, we conduct an extensive benchmarking study comparing existing state-of-the-art agentic systems such as ChatGPT Agent and chemical-specific data extraction agents. Additionally, we introduce our own single-agent approach that enables precise control over document preprocessing prior to extraction. We further evaluate the performance of modern baselines, such as GPT-5 and GPT-5 Thinking, to compare their capabilities with agentic approaches. Our empirical findings reveal persistent challenges in chemical information extraction, particularly in processing domain-specific terminology, complex tabular and schematic representations, and context-dependent ambiguities. The ChemX benchmark serves as a critical resource for advancing automated information extraction in chemistry, challenging the generalization capabilities of existing methods, and providing valuable insights into effective evaluation strategies.
format Preprint
id arxiv_https___arxiv_org_abs_2510_00795
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Benchmarking Agentic Systems in Automated Scientific Information Extraction with ChemX
Vepreva, Anastasia
Razlivina, Julia
Eremeeva, Maria
Gubina, Nina
Orlova, Anastasia
Dmitrenko, Aleksei
Kapranova, Ksenya
Jyakhwo, Susan
Vasilev, Nikita
Sarkisyan, Arsen
Chernyshov, Ivan Yu.
Vinogradov, Vladimir
Dmitrenko, Andrei
Artificial Intelligence
The emergence of agent-based systems represents a significant advancement in artificial intelligence, with growing applications in automated data extraction. However, chemical information extraction remains a formidable challenge due to the inherent heterogeneity of chemical data. Current agent-based approaches, both general-purpose and domain-specific, exhibit limited performance in this domain. To address this gap, we present ChemX, a comprehensive collection of 10 manually curated and domain-expert-validated datasets focusing on nanomaterials and small molecules. These datasets are designed to rigorously evaluate and enhance automated extraction methodologies in chemistry. To demonstrate their utility, we conduct an extensive benchmarking study comparing existing state-of-the-art agentic systems such as ChatGPT Agent and chemical-specific data extraction agents. Additionally, we introduce our own single-agent approach that enables precise control over document preprocessing prior to extraction. We further evaluate the performance of modern baselines, such as GPT-5 and GPT-5 Thinking, to compare their capabilities with agentic approaches. Our empirical findings reveal persistent challenges in chemical information extraction, particularly in processing domain-specific terminology, complex tabular and schematic representations, and context-dependent ambiguities. The ChemX benchmark serves as a critical resource for advancing automated information extraction in chemistry, challenging the generalization capabilities of existing methods, and providing valuable insights into effective evaluation strategies.
title Benchmarking Agentic Systems in Automated Scientific Information Extraction with ChemX
topic Artificial Intelligence
url https://arxiv.org/abs/2510.00795