MolParser: End-to-end Visual Recognition of Molecule Structures in the Wild

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Fang, Xi, Wang, Jiankun, Cai, Xiaochen, Chen, Shangqian, Yang, Shuwen, Tao, Haoyi, Wang, Nan, Yao, Lin, Zhang, Linfeng, Ke, Guolin
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918143903924224
author Fang, Xi
Wang, Jiankun
Cai, Xiaochen
Chen, Shangqian
Yang, Shuwen
Tao, Haoyi
Wang, Nan
Yao, Lin
Zhang, Linfeng
Ke, Guolin
author_facet Fang, Xi
Wang, Jiankun
Cai, Xiaochen
Chen, Shangqian
Yang, Shuwen
Tao, Haoyi
Wang, Nan
Yao, Lin
Zhang, Linfeng
Ke, Guolin
contents In recent decades, chemistry publications and patents have increased rapidly. A significant portion of key information is embedded in molecular structure figures, complicating large-scale literature searches and limiting the application of large language models in fields such as biology, chemistry, and pharmaceuticals. The automatic extraction of precise chemical structures is of critical importance. However, the presence of numerous Markush structures in real-world documents, along with variations in molecular image quality, drawing styles, and noise, significantly limits the performance of existing optical chemical structure recognition (OCSR) methods. We present MolParser, a novel end-to-end OCSR method that efficiently and accurately recognizes chemical structures from real-world documents, including difficult Markush structure. We use a extended SMILES encoding rule to annotate our training dataset. Under this rule, we build MolParser-7M, the largest annotated molecular image dataset to our knowledge. While utilizing a large amount of synthetic data, we employed active learning methods to incorporate substantial in-the-wild data, specifically samples cropped from real patents and scientific literature, into the training process. We trained an end-to-end molecular image captioning model, MolParser, using a curriculum learning approach. MolParser significantly outperforms classical and learning-based methods across most scenarios, with potential for broader downstream applications. The dataset is publicly available in huggingface.
format Preprint
id arxiv_https___arxiv_org_abs_2411_11098
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MolParser: End-to-end Visual Recognition of Molecule Structures in the Wild
Fang, Xi
Wang, Jiankun
Cai, Xiaochen
Chen, Shangqian
Yang, Shuwen
Tao, Haoyi
Wang, Nan
Yao, Lin
Zhang, Linfeng
Ke, Guolin
Computer Vision and Pattern Recognition
In recent decades, chemistry publications and patents have increased rapidly. A significant portion of key information is embedded in molecular structure figures, complicating large-scale literature searches and limiting the application of large language models in fields such as biology, chemistry, and pharmaceuticals. The automatic extraction of precise chemical structures is of critical importance. However, the presence of numerous Markush structures in real-world documents, along with variations in molecular image quality, drawing styles, and noise, significantly limits the performance of existing optical chemical structure recognition (OCSR) methods. We present MolParser, a novel end-to-end OCSR method that efficiently and accurately recognizes chemical structures from real-world documents, including difficult Markush structure. We use a extended SMILES encoding rule to annotate our training dataset. Under this rule, we build MolParser-7M, the largest annotated molecular image dataset to our knowledge. While utilizing a large amount of synthetic data, we employed active learning methods to incorporate substantial in-the-wild data, specifically samples cropped from real patents and scientific literature, into the training process. We trained an end-to-end molecular image captioning model, MolParser, using a curriculum learning approach. MolParser significantly outperforms classical and learning-based methods across most scenarios, with potential for broader downstream applications. The dataset is publicly available in huggingface.
title MolParser: End-to-end Visual Recognition of Molecule Structures in the Wild
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2411.11098