Saved in:
Bibliographic Details
Main Authors: Wang, Junying, Zhang, Zicheng, Guo, Yijin, Wen, Farong, Shen, Ye, Liang, Yingji, Wu, Yalun, Li, Wenzhe, Li, Chunyi, Chen, Zijian, Jia, Qi, Zhai, Guangtao
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2507.16514
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918151240810496
author Wang, Junying
Zhang, Zicheng
Guo, Yijin
Wen, Farong
Shen, Ye
Liang, Yingji
Wu, Yalun
Li, Wenzhe
Li, Chunyi
Chen, Zijian
Jia, Qi
Zhai, Guangtao
author_facet Wang, Junying
Zhang, Zicheng
Guo, Yijin
Wen, Farong
Shen, Ye
Liang, Yingji
Wu, Yalun
Li, Wenzhe
Li, Chunyi
Chen, Zijian
Jia, Qi
Zhai, Guangtao
contents As foundation models grow rapidly in capability and deployment, evaluating their scientific understanding becomes increasingly critical. Existing science benchmarks have made progress towards broad Range, wide Reach, and high Rigor, yet they often face two major challenges: data leakage risks that compromise benchmarking validity, and evaluation inefficiency due to large-scale testing. To address these issues, we introduce the Ever-Evolving Science Exam (EESE), a dynamic benchmark designed to reliably assess scientific capabilities in foundation models. Our approach consists of two components: 1) a non-public EESE-Pool with over 100K expertly constructed science instances (question-answer pairs) across 5 disciplines and 500+ subfields, built through a multi-stage pipeline ensuring Range, Reach, and Rigor, 2) a periodically updated 500-instance subset EESE, sampled and validated to enable leakage-resilient, low-overhead evaluations. Experiments on 32 open- and closed-source models demonstrate that EESE effectively differentiates the strengths and weaknesses of models in scientific fields and cognitive dimensions. Overall, EESE provides a robust, scalable, and forward-compatible solution for science benchmark design, offering a realistic measure of how well foundation models handle science questions. The project page is at: https://github.com/aiben-ch/EESE.
format Preprint
id arxiv_https___arxiv_org_abs_2507_16514
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The Ever-Evolving Science Exam
Wang, Junying
Zhang, Zicheng
Guo, Yijin
Wen, Farong
Shen, Ye
Liang, Yingji
Wu, Yalun
Li, Wenzhe
Li, Chunyi
Chen, Zijian
Jia, Qi
Zhai, Guangtao
Computation and Language
Artificial Intelligence
As foundation models grow rapidly in capability and deployment, evaluating their scientific understanding becomes increasingly critical. Existing science benchmarks have made progress towards broad Range, wide Reach, and high Rigor, yet they often face two major challenges: data leakage risks that compromise benchmarking validity, and evaluation inefficiency due to large-scale testing. To address these issues, we introduce the Ever-Evolving Science Exam (EESE), a dynamic benchmark designed to reliably assess scientific capabilities in foundation models. Our approach consists of two components: 1) a non-public EESE-Pool with over 100K expertly constructed science instances (question-answer pairs) across 5 disciplines and 500+ subfields, built through a multi-stage pipeline ensuring Range, Reach, and Rigor, 2) a periodically updated 500-instance subset EESE, sampled and validated to enable leakage-resilient, low-overhead evaluations. Experiments on 32 open- and closed-source models demonstrate that EESE effectively differentiates the strengths and weaknesses of models in scientific fields and cognitive dimensions. Overall, EESE provides a robust, scalable, and forward-compatible solution for science benchmark design, offering a realistic measure of how well foundation models handle science questions. The project page is at: https://github.com/aiben-ch/EESE.
title The Ever-Evolving Science Exam
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2507.16514