SciCustom: A Framework for Custom Evaluation of Scientific Capabilities in Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gu, Yiyang, Yang, Junwei, Luo, Junyu, Yuan, Ye, Feng, Bin, Xia, Yingce, Xie, Shufang, Liu, Kaili, Wu, Bohan, Shi, Qi, Li, Haoran, Xiao, Beier, Xiao, Zhiping, Luo, Xiao, Zhang, Weizhi, Yu, Philip S., Liu, Zequn, Zhang, Ming
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909055754174464
author Gu, Yiyang
Yang, Junwei
Luo, Junyu
Yuan, Ye
Feng, Bin
Xia, Yingce
Xie, Shufang
Liu, Kaili
Wu, Bohan
Shi, Qi
Li, Haoran
Xiao, Beier
Xiao, Zhiping
Luo, Xiao
Zhang, Weizhi
Yu, Philip S.
Liu, Zequn
Zhang, Ming
author_facet Gu, Yiyang
Yang, Junwei
Luo, Junyu
Yuan, Ye
Feng, Bin
Xia, Yingce
Xie, Shufang
Liu, Kaili
Wu, Bohan
Shi, Qi
Li, Haoran
Xiao, Beier
Xiao, Zhiping
Luo, Xiao
Zhang, Weizhi
Yu, Philip S.
Liu, Zequn
Zhang, Ming
contents Large language models (LLMs) are increasingly applied to scientific research, yet existing evaluations often fail to reflect the fine-grained capabilities required in practice. Most benchmarks are manually curated or domain-generic, limiting scalability and alignment with real scientific use cases. In this paper, we propose a new framework named SciCustom to address the problem. It enables the custom construction of benchmarks from large-scale scientific data to evaluate application-specific scientific capabilities in LLMs. SciCustom first organizes scientific knowledge into ontology-grounded knowledge units with controlled granularity and trains a tagger to map large-scale data instances into this knowledge space. Given a custom requirement, relevant knowledge units are identified via voting-based multi-model consensus. These units enable relevance-aware benchmark retrieval via binary search, followed by proxy subset selection and data-grounded benchmark generation for efficient evaluation. Experiments in chemistry and healthcare demonstrate that SciCustom reveals fine-grained differences in LLM scientific capabilities that standard benchmarks overlook, while requiring neither expert annotation nor synthetic question generation. This work provides a scalable and application-aware foundation for benchmarking scientific capabilities in LLMs. The source code is available at https://github.com/yjwtheonly/SciCustom.
format Preprint
id arxiv_https___arxiv_org_abs_2605_19357
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SciCustom: A Framework for Custom Evaluation of Scientific Capabilities in Large Language Models
Gu, Yiyang
Yang, Junwei
Luo, Junyu
Yuan, Ye
Feng, Bin
Xia, Yingce
Xie, Shufang
Liu, Kaili
Wu, Bohan
Shi, Qi
Li, Haoran
Xiao, Beier
Xiao, Zhiping
Luo, Xiao
Zhang, Weizhi
Yu, Philip S.
Liu, Zequn
Zhang, Ming
Computation and Language
Large language models (LLMs) are increasingly applied to scientific research, yet existing evaluations often fail to reflect the fine-grained capabilities required in practice. Most benchmarks are manually curated or domain-generic, limiting scalability and alignment with real scientific use cases. In this paper, we propose a new framework named SciCustom to address the problem. It enables the custom construction of benchmarks from large-scale scientific data to evaluate application-specific scientific capabilities in LLMs. SciCustom first organizes scientific knowledge into ontology-grounded knowledge units with controlled granularity and trains a tagger to map large-scale data instances into this knowledge space. Given a custom requirement, relevant knowledge units are identified via voting-based multi-model consensus. These units enable relevance-aware benchmark retrieval via binary search, followed by proxy subset selection and data-grounded benchmark generation for efficient evaluation. Experiments in chemistry and healthcare demonstrate that SciCustom reveals fine-grained differences in LLM scientific capabilities that standard benchmarks overlook, while requiring neither expert annotation nor synthetic question generation. This work provides a scalable and application-aware foundation for benchmarking scientific capabilities in LLMs. The source code is available at https://github.com/yjwtheonly/SciCustom.
title SciCustom: A Framework for Custom Evaluation of Scientific Capabilities in Large Language Models
topic Computation and Language
url https://arxiv.org/abs/2605.19357