ChemEval: A Comprehensive Multi-Level Chemical Evaluation for Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Yuqing, Zhang, Rongyang, He, Xuesong, Zhi, Xuyang, Wang, Hao, Li, Xin, Xu, Feiyang, Liu, Deguang, Liang, Huadong, Li, Yi, Cui, Jian, Liu, Zimu, Wang, Shijin, Hu, Guoping, Liu, Guiquan, Liu, Qi, Lian, Defu, Chen, Enhong
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910615957667840
author Huang, Yuqing
Zhang, Rongyang
He, Xuesong
Zhi, Xuyang
Wang, Hao
Li, Xin
Xu, Feiyang
Liu, Deguang
Liang, Huadong
Li, Yi
Cui, Jian
Liu, Zimu
Wang, Shijin
Hu, Guoping
Liu, Guiquan
Liu, Qi
Lian, Defu
Chen, Enhong
author_facet Huang, Yuqing
Zhang, Rongyang
He, Xuesong
Zhi, Xuyang
Wang, Hao
Li, Xin
Xu, Feiyang
Liu, Deguang
Liang, Huadong
Li, Yi
Cui, Jian
Liu, Zimu
Wang, Shijin
Hu, Guoping
Liu, Guiquan
Liu, Qi
Lian, Defu
Chen, Enhong
contents There is a growing interest in the role that LLMs play in chemistry which lead to an increased focus on the development of LLMs benchmarks tailored to chemical domains to assess the performance of LLMs across a spectrum of chemical tasks varying in type and complexity. However, existing benchmarks in this domain fail to adequately meet the specific requirements of chemical research professionals. To this end, we propose \textbf{\textit{ChemEval}}, which provides a comprehensive assessment of the capabilities of LLMs across a wide range of chemical domain tasks. Specifically, ChemEval identified 4 crucial progressive levels in chemistry, assessing 12 dimensions of LLMs across 42 distinct chemical tasks which are informed by open-source data and the data meticulously crafted by chemical experts, ensuring that the tasks have practical value and can effectively evaluate the capabilities of LLMs. In the experiment, we evaluate 12 mainstream LLMs on ChemEval under zero-shot and few-shot learning contexts, which included carefully selected demonstration examples and carefully designed prompts. The results show that while general LLMs like GPT-4 and Claude-3.5 excel in literature understanding and instruction following, they fall short in tasks demanding advanced chemical knowledge. Conversely, specialized LLMs exhibit enhanced chemical competencies, albeit with reduced literary comprehension. This suggests that LLMs have significant potential for enhancement when tackling sophisticated tasks in the field of chemistry. We believe our work will facilitate the exploration of their potential to drive progress in chemistry. Our benchmark and analysis will be available at {\color{blue} \url{https://github.com/USTC-StarTeam/ChemEval}}.
format Preprint
id arxiv_https___arxiv_org_abs_2409_13989
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ChemEval: A Comprehensive Multi-Level Chemical Evaluation for Large Language Models
Huang, Yuqing
Zhang, Rongyang
He, Xuesong
Zhi, Xuyang
Wang, Hao
Li, Xin
Xu, Feiyang
Liu, Deguang
Liang, Huadong
Li, Yi
Cui, Jian
Liu, Zimu
Wang, Shijin
Hu, Guoping
Liu, Guiquan
Liu, Qi
Lian, Defu
Chen, Enhong
Computation and Language
Artificial Intelligence
Machine Learning
Chemical Physics
Biomolecules
There is a growing interest in the role that LLMs play in chemistry which lead to an increased focus on the development of LLMs benchmarks tailored to chemical domains to assess the performance of LLMs across a spectrum of chemical tasks varying in type and complexity. However, existing benchmarks in this domain fail to adequately meet the specific requirements of chemical research professionals. To this end, we propose \textbf{\textit{ChemEval}}, which provides a comprehensive assessment of the capabilities of LLMs across a wide range of chemical domain tasks. Specifically, ChemEval identified 4 crucial progressive levels in chemistry, assessing 12 dimensions of LLMs across 42 distinct chemical tasks which are informed by open-source data and the data meticulously crafted by chemical experts, ensuring that the tasks have practical value and can effectively evaluate the capabilities of LLMs. In the experiment, we evaluate 12 mainstream LLMs on ChemEval under zero-shot and few-shot learning contexts, which included carefully selected demonstration examples and carefully designed prompts. The results show that while general LLMs like GPT-4 and Claude-3.5 excel in literature understanding and instruction following, they fall short in tasks demanding advanced chemical knowledge. Conversely, specialized LLMs exhibit enhanced chemical competencies, albeit with reduced literary comprehension. This suggests that LLMs have significant potential for enhancement when tackling sophisticated tasks in the field of chemistry. We believe our work will facilitate the exploration of their potential to drive progress in chemistry. Our benchmark and analysis will be available at {\color{blue} \url{https://github.com/USTC-StarTeam/ChemEval}}.
title ChemEval: A Comprehensive Multi-Level Chemical Evaluation for Large Language Models
topic Computation and Language
Artificial Intelligence
Machine Learning
Chemical Physics
Biomolecules
url https://arxiv.org/abs/2409.13989