zERExtractor:An Automated Platform for Enzyme-Catalyzed Reaction Data Extraction from Scientific Literature

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Zhou, Rui, Ma, Haohui, Xin, Tianle, Zou, Lixin, Hu, Qiuyue, Cheng, Hongxi, Lin, Mingzhi, Guo, Jingjing, Wang, Sheng, Zhang, Guoqing, Wei, Yanjie, Zheng, Liangzhen
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912537106186240
author Zhou, Rui
Ma, Haohui
Xin, Tianle
Zou, Lixin
Hu, Qiuyue
Cheng, Hongxi
Lin, Mingzhi
Guo, Jingjing
Wang, Sheng
Zhang, Guoqing
Wei, Yanjie
Zheng, Liangzhen
author_facet Zhou, Rui
Ma, Haohui
Xin, Tianle
Zou, Lixin
Hu, Qiuyue
Cheng, Hongxi
Lin, Mingzhi
Guo, Jingjing
Wang, Sheng
Zhang, Guoqing
Wei, Yanjie
Zheng, Liangzhen
contents The rapid expansion of enzyme kinetics literature has outpaced the curation capabilities of major biochemical databases, creating a substantial barrier to AI-driven modeling and knowledge discovery. We present zERExtractor, an automated and extensible platform for comprehensive extraction of enzyme-catalyzed reaction and activity data from scientific literature. zERExtractor features a unified, modular architecture that supports plug-and-play integration of state-of-the-art models, including large language models (LLMs), as interchangeable components, enabling continuous system evolution alongside advances in AI. Our pipeline combines domain-adapted deep learning, advanced OCR, semantic entity recognition, and prompt-driven LLM modules, together with human expert corrections, to extract kinetic parameters (e.g., kcat, Km), enzyme sequences, substrate SMILES, experimental conditions, and molecular diagrams from heterogeneous document formats. Through active learning strategies integrating AI-assisted annotation, expert validation, and iterative refinement, the system adapts rapidly to new data sources. We also release a large benchmark dataset comprising over 1,000 annotated tables and 5,000 biological fields from 270 P450-related enzymology publications. Benchmarking demonstrates that zERExtractor consistently outperforms existing baselines in table recognition (Acc 89.9%), molecular image interpretation (up to 99.1%), and relation extraction (accuracy 94.2%). zERExtractor bridges the longstanding data gap in enzyme kinetics with a flexible, plugin-ready framework and high-fidelity extraction, laying the groundwork for future AI-powered enzyme modeling and biochemical knowledge discovery.
format Preprint
id arxiv_https___arxiv_org_abs_2508_09995
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle zERExtractor:An Automated Platform for Enzyme-Catalyzed Reaction Data Extraction from Scientific Literature
Zhou, Rui
Ma, Haohui
Xin, Tianle
Zou, Lixin
Hu, Qiuyue
Cheng, Hongxi
Lin, Mingzhi
Guo, Jingjing
Wang, Sheng
Zhang, Guoqing
Wei, Yanjie
Zheng, Liangzhen
Biomolecules
Emerging Technologies
Machine Learning
The rapid expansion of enzyme kinetics literature has outpaced the curation capabilities of major biochemical databases, creating a substantial barrier to AI-driven modeling and knowledge discovery. We present zERExtractor, an automated and extensible platform for comprehensive extraction of enzyme-catalyzed reaction and activity data from scientific literature. zERExtractor features a unified, modular architecture that supports plug-and-play integration of state-of-the-art models, including large language models (LLMs), as interchangeable components, enabling continuous system evolution alongside advances in AI. Our pipeline combines domain-adapted deep learning, advanced OCR, semantic entity recognition, and prompt-driven LLM modules, together with human expert corrections, to extract kinetic parameters (e.g., kcat, Km), enzyme sequences, substrate SMILES, experimental conditions, and molecular diagrams from heterogeneous document formats. Through active learning strategies integrating AI-assisted annotation, expert validation, and iterative refinement, the system adapts rapidly to new data sources. We also release a large benchmark dataset comprising over 1,000 annotated tables and 5,000 biological fields from 270 P450-related enzymology publications. Benchmarking demonstrates that zERExtractor consistently outperforms existing baselines in table recognition (Acc 89.9%), molecular image interpretation (up to 99.1%), and relation extraction (accuracy 94.2%). zERExtractor bridges the longstanding data gap in enzyme kinetics with a flexible, plugin-ready framework and high-fidelity extraction, laying the groundwork for future AI-powered enzyme modeling and biochemical knowledge discovery.
title zERExtractor:An Automated Platform for Enzyme-Catalyzed Reaction Data Extraction from Scientific Literature
topic Biomolecules
Emerging Technologies
Machine Learning
url https://arxiv.org/abs/2508.09995