SciHorizon: Benchmarking AI-for-Science Readiness from Scientific Data to Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Qin, Chuan, Chen, Xin, Wang, Chengrui, Wu, Pengmin, Chen, Xi, Cheng, Yihang, Zhao, Jingyi, Xiao, Meng, Dong, Xiangchao, Long, Qingqing, Pan, Boya, Wu, Han, Li, Chengzan, Zhou, Yuanchun, Xiong, Hui, Zhu, Hengshu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912400455761920
author Qin, Chuan
Chen, Xin
Wang, Chengrui
Wu, Pengmin
Chen, Xi
Cheng, Yihang
Zhao, Jingyi
Xiao, Meng
Dong, Xiangchao
Long, Qingqing
Pan, Boya
Wu, Han
Li, Chengzan
Zhou, Yuanchun
Xiong, Hui
Zhu, Hengshu
author_facet Qin, Chuan
Chen, Xin
Wang, Chengrui
Wu, Pengmin
Chen, Xi
Cheng, Yihang
Zhao, Jingyi
Xiao, Meng
Dong, Xiangchao
Long, Qingqing
Pan, Boya
Wu, Han
Li, Chengzan
Zhou, Yuanchun
Xiong, Hui
Zhu, Hengshu
contents In recent years, the rapid advancement of Artificial Intelligence (AI) technologies, particularly Large Language Models (LLMs), has revolutionized the paradigm of scientific discovery, establishing AI-for-Science (AI4Science) as a dynamic and evolving field. However, there is still a lack of an effective framework for the overall assessment of AI4Science, particularly from a holistic perspective on data quality and model capability. Therefore, in this study, we propose SciHorizon, a comprehensive assessment framework designed to benchmark the readiness of AI4Science from both scientific data and LLM perspectives. First, we introduce a generalizable framework for assessing AI-ready scientific data, encompassing four key dimensions: Quality, FAIRness, Explainability, and Compliance-which are subdivided into 15 sub-dimensions. Drawing on data resource papers published between 2018 and 2023 in peer-reviewed journals, we present recommendation lists of AI-ready datasets for Earth, Life, and Materials Sciences, making a novel and original contribution to the field. Concurrently, to assess the capabilities of LLMs across multiple scientific disciplines, we establish 16 assessment dimensions based on five core indicators Knowledge, Understanding, Reasoning, Multimodality, and Values spanning Mathematics, Physics, Chemistry, Life Sciences, and Earth and Space Sciences. Using the developed benchmark datasets, we have conducted a comprehensive evaluation of over 50 representative open-source and closed source LLMs. All the results are publicly available and can be accessed online at www.scihorizon.cn/en.
format Preprint
id arxiv_https___arxiv_org_abs_2503_13503
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SciHorizon: Benchmarking AI-for-Science Readiness from Scientific Data to Large Language Models
Qin, Chuan
Chen, Xin
Wang, Chengrui
Wu, Pengmin
Chen, Xi
Cheng, Yihang
Zhao, Jingyi
Xiao, Meng
Dong, Xiangchao
Long, Qingqing
Pan, Boya
Wu, Han
Li, Chengzan
Zhou, Yuanchun
Xiong, Hui
Zhu, Hengshu
Machine Learning
Computation and Language
Digital Libraries
Information Retrieval
In recent years, the rapid advancement of Artificial Intelligence (AI) technologies, particularly Large Language Models (LLMs), has revolutionized the paradigm of scientific discovery, establishing AI-for-Science (AI4Science) as a dynamic and evolving field. However, there is still a lack of an effective framework for the overall assessment of AI4Science, particularly from a holistic perspective on data quality and model capability. Therefore, in this study, we propose SciHorizon, a comprehensive assessment framework designed to benchmark the readiness of AI4Science from both scientific data and LLM perspectives. First, we introduce a generalizable framework for assessing AI-ready scientific data, encompassing four key dimensions: Quality, FAIRness, Explainability, and Compliance-which are subdivided into 15 sub-dimensions. Drawing on data resource papers published between 2018 and 2023 in peer-reviewed journals, we present recommendation lists of AI-ready datasets for Earth, Life, and Materials Sciences, making a novel and original contribution to the field. Concurrently, to assess the capabilities of LLMs across multiple scientific disciplines, we establish 16 assessment dimensions based on five core indicators Knowledge, Understanding, Reasoning, Multimodality, and Values spanning Mathematics, Physics, Chemistry, Life Sciences, and Earth and Space Sciences. Using the developed benchmark datasets, we have conducted a comprehensive evaluation of over 50 representative open-source and closed source LLMs. All the results are publicly available and can be accessed online at www.scihorizon.cn/en.
title SciHorizon: Benchmarking AI-for-Science Readiness from Scientific Data to Large Language Models
topic Machine Learning
Computation and Language
Digital Libraries
Information Retrieval
url https://arxiv.org/abs/2503.13503