SUPERChem: A Multimodal Reasoning Benchmark in Chemistry

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhao, Zehua, Huang, Zhixian, Li, Junren, Lin, Siyu, Zhou, Junting, Cao, Fengqi, Zhou, Kun, Ge, Rui, Long, Tingting, Zhu, Yuexiang, Liu, Yan, Zheng, Jie, Wei, Junnian, Zhu, Rong, Zou, Peng, Li, Wenyu, Cheng, Zekai, Ding, Tian, Wang, Yaxuan, Yan, Yizhao, Wei, Tingru, Ming, Haowei, Mao, Weijie, Sun, Chen, Liu, Yiming, Wang, Zichen, Zhang, Zuo, Yang, Tong, Ma, Hao, Gao, Zhen, Pei, Jian
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909937520607232
author Zhao, Zehua
Huang, Zhixian
Li, Junren
Lin, Siyu
Zhou, Junting
Cao, Fengqi
Zhou, Kun
Ge, Rui
Long, Tingting
Zhu, Yuexiang
Liu, Yan
Zheng, Jie
Wei, Junnian
Zhu, Rong
Zou, Peng
Li, Wenyu
Cheng, Zekai
Ding, Tian
Wang, Yaxuan
Yan, Yizhao
Wei, Tingru
Ming, Haowei
Mao, Weijie
Sun, Chen
Liu, Yiming
Wang, Zichen
Zhang, Zuo
Yang, Tong
Ma, Hao
Gao, Zhen
Pei, Jian
author_facet Zhao, Zehua
Huang, Zhixian
Li, Junren
Lin, Siyu
Zhou, Junting
Cao, Fengqi
Zhou, Kun
Ge, Rui
Long, Tingting
Zhu, Yuexiang
Liu, Yan
Zheng, Jie
Wei, Junnian
Zhu, Rong
Zou, Peng
Li, Wenyu
Cheng, Zekai
Ding, Tian
Wang, Yaxuan
Yan, Yizhao
Wei, Tingru
Ming, Haowei
Mao, Weijie
Sun, Chen
Liu, Yiming
Wang, Zichen
Zhang, Zuo
Yang, Tong
Ma, Hao
Gao, Zhen
Pei, Jian
contents Current benchmarks for evaluating the chemical reasoning capabilities of Large Language Models (LLMs) are limited by oversimplified tasks, lack of process-level evaluation, and misalignment with expert-level chemistry skills. To address these issues, we introduce SUPERChem, a benchmark of 500 expert-curated reasoning-intensive chemistry problems, covering diverse subfields and provided in both multimodal and text-only formats. Original content and an iterative curation pipeline eliminate flawed items and mitigate data contamination. Each problem is paired with an expert-authored solution path, enabling Reasoning Path Fidelity (RPF) scoring to evaluate reasoning quality beyond final-answer accuracy. Evaluations against a human baseline of 40.3% accuracy show that even the best-performing model, GPT-5 (High), reaches only 38.5%, followed closely by Gemini 2.5 Pro (37.9%) and DeepSeek-V3.1-Think (37.3%). SUPERChem elicits multi-step, multimodal reasoning, reveals model-dependent effects of visual information, and distinguishes high-fidelity reasoners from heuristic ones. By providing a challenging benchmark and a reliable evaluation framework, SUPERChem aims to facilitate the advancement of LLMs toward expert-level chemical intelligence. The dataset of the benchmark is available at https://huggingface.co/datasets/ZehuaZhao/SUPERChem.
format Preprint
id arxiv_https___arxiv_org_abs_2512_01274
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SUPERChem: A Multimodal Reasoning Benchmark in Chemistry
Zhao, Zehua
Huang, Zhixian
Li, Junren
Lin, Siyu
Zhou, Junting
Cao, Fengqi
Zhou, Kun
Ge, Rui
Long, Tingting
Zhu, Yuexiang
Liu, Yan
Zheng, Jie
Wei, Junnian
Zhu, Rong
Zou, Peng
Li, Wenyu
Cheng, Zekai
Ding, Tian
Wang, Yaxuan
Yan, Yizhao
Wei, Tingru
Ming, Haowei
Mao, Weijie
Sun, Chen
Liu, Yiming
Wang, Zichen
Zhang, Zuo
Yang, Tong
Ma, Hao
Gao, Zhen
Pei, Jian
Computation and Language
Artificial Intelligence
Machine Learning
Current benchmarks for evaluating the chemical reasoning capabilities of Large Language Models (LLMs) are limited by oversimplified tasks, lack of process-level evaluation, and misalignment with expert-level chemistry skills. To address these issues, we introduce SUPERChem, a benchmark of 500 expert-curated reasoning-intensive chemistry problems, covering diverse subfields and provided in both multimodal and text-only formats. Original content and an iterative curation pipeline eliminate flawed items and mitigate data contamination. Each problem is paired with an expert-authored solution path, enabling Reasoning Path Fidelity (RPF) scoring to evaluate reasoning quality beyond final-answer accuracy. Evaluations against a human baseline of 40.3% accuracy show that even the best-performing model, GPT-5 (High), reaches only 38.5%, followed closely by Gemini 2.5 Pro (37.9%) and DeepSeek-V3.1-Think (37.3%). SUPERChem elicits multi-step, multimodal reasoning, reveals model-dependent effects of visual information, and distinguishes high-fidelity reasoners from heuristic ones. By providing a challenging benchmark and a reliable evaluation framework, SUPERChem aims to facilitate the advancement of LLMs toward expert-level chemical intelligence. The dataset of the benchmark is available at https://huggingface.co/datasets/ZehuaZhao/SUPERChem.
title SUPERChem: A Multimodal Reasoning Benchmark in Chemistry
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2512.01274