ChemVA: Advancing Large Language Models on Chemical Reaction Diagrams Understanding

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Rao, Mingyang, Feng, Kehua, Zhu, Zhihui, Fu, Jiangzhen, Yu, Hao, Ding, Keyan, Chen, Huajun
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917505193213952
author Rao, Mingyang
Feng, Kehua
Zhu, Zhihui
Fu, Jiangzhen
Yu, Hao
Ding, Keyan
Chen, Huajun
author_facet Rao, Mingyang
Feng, Kehua
Zhu, Zhihui
Fu, Jiangzhen
Yu, Hao
Ding, Keyan
Chen, Huajun
contents While Large Language Models (LLMs) have revolutionized scientific text processing, they exhibit a significant capability gap when interpreting chemical reaction diagrams. We identify two fundamental bottlenecks restricting current systems: a Visual Deficit, where generic vision encoders struggle to resolve the strict topological connectivity of dense molecular graphs, and a Semantic Disconnect, where standard linear strings, such as SMILES, fail to effectively activate the model's latent chemical reasoning. To bridge these gaps, we propose the Chemical Visual Activation (ChemVA) framework, which employs a Visual Anchor mechanism to ground functional groups via hybrid-granularity detection, followed by a semantic alignment approach that translates visual features into entity names to maximize knowledge activation in LLMs. We evaluate our approach on OCRD-Bench, a newly constructed dataset featuring dense visual-semantic contexts and comprehensive reaction coverage to evaluate the full spectrum from recognition to reasoning. Extensive experiments on OCRD-Bench demonstrate that ChemVA achieves 92.0% structural recognition accuracy. By bridging visual and semantic bottlenecks, our framework delivers a consistent performance gain of approximately 20 percentage points across 9 diverse LLMs, enabling open-weight models to rival proprietary SOTA systems in complex chemical reasoning tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2605_17214
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ChemVA: Advancing Large Language Models on Chemical Reaction Diagrams Understanding
Rao, Mingyang
Feng, Kehua
Zhu, Zhihui
Fu, Jiangzhen
Yu, Hao
Ding, Keyan
Chen, Huajun
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
While Large Language Models (LLMs) have revolutionized scientific text processing, they exhibit a significant capability gap when interpreting chemical reaction diagrams. We identify two fundamental bottlenecks restricting current systems: a Visual Deficit, where generic vision encoders struggle to resolve the strict topological connectivity of dense molecular graphs, and a Semantic Disconnect, where standard linear strings, such as SMILES, fail to effectively activate the model's latent chemical reasoning. To bridge these gaps, we propose the Chemical Visual Activation (ChemVA) framework, which employs a Visual Anchor mechanism to ground functional groups via hybrid-granularity detection, followed by a semantic alignment approach that translates visual features into entity names to maximize knowledge activation in LLMs. We evaluate our approach on OCRD-Bench, a newly constructed dataset featuring dense visual-semantic contexts and comprehensive reaction coverage to evaluate the full spectrum from recognition to reasoning. Extensive experiments on OCRD-Bench demonstrate that ChemVA achieves 92.0% structural recognition accuracy. By bridging visual and semantic bottlenecks, our framework delivers a consistent performance gain of approximately 20 percentage points across 9 diverse LLMs, enabling open-weight models to rival proprietary SOTA systems in complex chemical reasoning tasks.
title ChemVA: Advancing Large Language Models on Chemical Reaction Diagrams Understanding
topic Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.17214