SUPERChem: A Multimodal Reasoning Benchmark in Chemistry
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909937520607232 |
|---|---|
| author | Zhao, Zehua Huang, Zhixian Li, Junren Lin, Siyu Zhou, Junting Cao, Fengqi Zhou, Kun Ge, Rui Long, Tingting Zhu, Yuexiang Liu, Yan Zheng, Jie Wei, Junnian Zhu, Rong Zou, Peng Li, Wenyu Cheng, Zekai Ding, Tian Wang, Yaxuan Yan, Yizhao Wei, Tingru Ming, Haowei Mao, Weijie Sun, Chen Liu, Yiming Wang, Zichen Zhang, Zuo Yang, Tong Ma, Hao Gao, Zhen Pei, Jian |
| author_facet | Zhao, Zehua Huang, Zhixian Li, Junren Lin, Siyu Zhou, Junting Cao, Fengqi Zhou, Kun Ge, Rui Long, Tingting Zhu, Yuexiang Liu, Yan Zheng, Jie Wei, Junnian Zhu, Rong Zou, Peng Li, Wenyu Cheng, Zekai Ding, Tian Wang, Yaxuan Yan, Yizhao Wei, Tingru Ming, Haowei Mao, Weijie Sun, Chen Liu, Yiming Wang, Zichen Zhang, Zuo Yang, Tong Ma, Hao Gao, Zhen Pei, Jian |
| contents | Current benchmarks for evaluating the chemical reasoning capabilities of Large Language Models (LLMs) are limited by oversimplified tasks, lack of process-level evaluation, and misalignment with expert-level chemistry skills. To address these issues, we introduce SUPERChem, a benchmark of 500 expert-curated reasoning-intensive chemistry problems, covering diverse subfields and provided in both multimodal and text-only formats. Original content and an iterative curation pipeline eliminate flawed items and mitigate data contamination. Each problem is paired with an expert-authored solution path, enabling Reasoning Path Fidelity (RPF) scoring to evaluate reasoning quality beyond final-answer accuracy. Evaluations against a human baseline of 40.3% accuracy show that even the best-performing model, GPT-5 (High), reaches only 38.5%, followed closely by Gemini 2.5 Pro (37.9%) and DeepSeek-V3.1-Think (37.3%). SUPERChem elicits multi-step, multimodal reasoning, reveals model-dependent effects of visual information, and distinguishes high-fidelity reasoners from heuristic ones. By providing a challenging benchmark and a reliable evaluation framework, SUPERChem aims to facilitate the advancement of LLMs toward expert-level chemical intelligence. The dataset of the benchmark is available at https://huggingface.co/datasets/ZehuaZhao/SUPERChem. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2512_01274 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | SUPERChem: A Multimodal Reasoning Benchmark in Chemistry Zhao, Zehua Huang, Zhixian Li, Junren Lin, Siyu Zhou, Junting Cao, Fengqi Zhou, Kun Ge, Rui Long, Tingting Zhu, Yuexiang Liu, Yan Zheng, Jie Wei, Junnian Zhu, Rong Zou, Peng Li, Wenyu Cheng, Zekai Ding, Tian Wang, Yaxuan Yan, Yizhao Wei, Tingru Ming, Haowei Mao, Weijie Sun, Chen Liu, Yiming Wang, Zichen Zhang, Zuo Yang, Tong Ma, Hao Gao, Zhen Pei, Jian Computation and Language Artificial Intelligence Machine Learning Current benchmarks for evaluating the chemical reasoning capabilities of Large Language Models (LLMs) are limited by oversimplified tasks, lack of process-level evaluation, and misalignment with expert-level chemistry skills. To address these issues, we introduce SUPERChem, a benchmark of 500 expert-curated reasoning-intensive chemistry problems, covering diverse subfields and provided in both multimodal and text-only formats. Original content and an iterative curation pipeline eliminate flawed items and mitigate data contamination. Each problem is paired with an expert-authored solution path, enabling Reasoning Path Fidelity (RPF) scoring to evaluate reasoning quality beyond final-answer accuracy. Evaluations against a human baseline of 40.3% accuracy show that even the best-performing model, GPT-5 (High), reaches only 38.5%, followed closely by Gemini 2.5 Pro (37.9%) and DeepSeek-V3.1-Think (37.3%). SUPERChem elicits multi-step, multimodal reasoning, reveals model-dependent effects of visual information, and distinguishes high-fidelity reasoners from heuristic ones. By providing a challenging benchmark and a reliable evaluation framework, SUPERChem aims to facilitate the advancement of LLMs toward expert-level chemical intelligence. The dataset of the benchmark is available at https://huggingface.co/datasets/ZehuaZhao/SUPERChem. |
| title | SUPERChem: A Multimodal Reasoning Benchmark in Chemistry |
| topic | Computation and Language Artificial Intelligence Machine Learning |
| url | https://arxiv.org/abs/2512.01274 |