RxnBench: A Multimodal Benchmark for Evaluating Large Language Models on Chemical Reaction Understanding from Scientific Literature

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Li, Hanzheng, Fang, Xi, Li, Yixuan, Huang, Chaozheng, Wang, Junjie, Wang, Xi, Bai, Hongzhe, Hao, Bojun, Lin, Shenyu, Liang, Huiqi, Zhang, Linfeng, Ke, Guolin
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908792968445952
author Li, Hanzheng
Fang, Xi
Li, Yixuan
Huang, Chaozheng
Wang, Junjie
Wang, Xi
Bai, Hongzhe
Hao, Bojun
Lin, Shenyu
Liang, Huiqi
Zhang, Linfeng
Ke, Guolin
author_facet Li, Hanzheng
Fang, Xi
Li, Yixuan
Huang, Chaozheng
Wang, Junjie
Wang, Xi
Bai, Hongzhe
Hao, Bojun
Lin, Shenyu
Liang, Huiqi
Zhang, Linfeng
Ke, Guolin
contents The integration of Multimodal Large Language Models (MLLMs) into chemistry promises to revolutionize scientific discovery, yet their ability to comprehend the dense, graphical language of reactions within authentic literature remains underexplored. Here, we introduce RxnBench, a multi-tiered benchmark designed to rigorously evaluate MLLMs on chemical reaction understanding from scientific PDFs. RxnBench comprises two tasks: Single-Figure QA (SF-QA), which tests fine-grained visual perception and mechanistic reasoning using 1,525 questions derived from 305 curated reaction schemes, and Full-Document QA (FD-QA), which challenges models to synthesize information from 108 articles, requiring cross-modal integration of text, schemes, and tables. Our evaluation of MLLMs reveals a critical capability gap: while models excel at extracting explicit text, they struggle with deep chemical logic and precise structural recognition. Notably, models with inference-time reasoning significantly outperform standard architectures, yet none achieve 50\% accuracy on FD-QA. These findings underscore the urgent need for domain-specific visual encoders and stronger reasoning engines to advance autonomous AI chemists.
format Preprint
id arxiv_https___arxiv_org_abs_2512_23565
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RxnBench: A Multimodal Benchmark for Evaluating Large Language Models on Chemical Reaction Understanding from Scientific Literature
Li, Hanzheng
Fang, Xi
Li, Yixuan
Huang, Chaozheng
Wang, Junjie
Wang, Xi
Bai, Hongzhe
Hao, Bojun
Lin, Shenyu
Liang, Huiqi
Zhang, Linfeng
Ke, Guolin
Computer Vision and Pattern Recognition
Artificial Intelligence
The integration of Multimodal Large Language Models (MLLMs) into chemistry promises to revolutionize scientific discovery, yet their ability to comprehend the dense, graphical language of reactions within authentic literature remains underexplored. Here, we introduce RxnBench, a multi-tiered benchmark designed to rigorously evaluate MLLMs on chemical reaction understanding from scientific PDFs. RxnBench comprises two tasks: Single-Figure QA (SF-QA), which tests fine-grained visual perception and mechanistic reasoning using 1,525 questions derived from 305 curated reaction schemes, and Full-Document QA (FD-QA), which challenges models to synthesize information from 108 articles, requiring cross-modal integration of text, schemes, and tables. Our evaluation of MLLMs reveals a critical capability gap: while models excel at extracting explicit text, they struggle with deep chemical logic and precise structural recognition. Notably, models with inference-time reasoning significantly outperform standard architectures, yet none achieve 50\% accuracy on FD-QA. These findings underscore the urgent need for domain-specific visual encoders and stronger reasoning engines to advance autonomous AI chemists.
title RxnBench: A Multimodal Benchmark for Evaluating Large Language Models on Chemical Reaction Understanding from Scientific Literature
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2512.23565