Saved in:
Bibliographic Details
Main Authors: Tian, Minyang, Gao, Luyu, Zhang, Shizhuo Dylan, Chen, Xinan, Fan, Cunwei, Guo, Xuefei, Haas, Roland, Ji, Pan, Krongchon, Kittithat, Li, Yao, Liu, Shengyan, Luo, Di, Ma, Yutao, Tong, Hao, Trinh, Kha, Tian, Chenyu, Wang, Zihan, Wu, Bohao, Xiong, Yanyu, Yin, Shengzhu, Zhu, Minhui, Lieret, Kilian, Lu, Yanxin, Liu, Genglin, Du, Yufeng, Tao, Tianhua, Press, Ofir, Callan, Jamie, Huerta, Eliu, Peng, Hao
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2407.13168
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911960581275648
author Tian, Minyang
Gao, Luyu
Zhang, Shizhuo Dylan
Chen, Xinan
Fan, Cunwei
Guo, Xuefei
Haas, Roland
Ji, Pan
Krongchon, Kittithat
Li, Yao
Liu, Shengyan
Luo, Di
Ma, Yutao
Tong, Hao
Trinh, Kha
Tian, Chenyu
Wang, Zihan
Wu, Bohao
Xiong, Yanyu
Yin, Shengzhu
Zhu, Minhui
Lieret, Kilian
Lu, Yanxin
Liu, Genglin
Du, Yufeng
Tao, Tianhua
Press, Ofir
Callan, Jamie
Huerta, Eliu
Peng, Hao
author_facet Tian, Minyang
Gao, Luyu
Zhang, Shizhuo Dylan
Chen, Xinan
Fan, Cunwei
Guo, Xuefei
Haas, Roland
Ji, Pan
Krongchon, Kittithat
Li, Yao
Liu, Shengyan
Luo, Di
Ma, Yutao
Tong, Hao
Trinh, Kha
Tian, Chenyu
Wang, Zihan
Wu, Bohao
Xiong, Yanyu
Yin, Shengzhu
Zhu, Minhui
Lieret, Kilian
Lu, Yanxin
Liu, Genglin
Du, Yufeng
Tao, Tianhua
Press, Ofir
Callan, Jamie
Huerta, Eliu
Peng, Hao
contents Since language models (LMs) now outperform average humans on many challenging tasks, it has become increasingly difficult to develop challenging, high-quality, and realistic evaluations. We address this issue by examining LMs' capabilities to generate code for solving real scientific research problems. Incorporating input from scientists and AI researchers in 16 diverse natural science sub-fields, including mathematics, physics, chemistry, biology, and materials science, we created a scientist-curated coding benchmark, SciCode. The problems in SciCode naturally factorize into multiple subproblems, each involving knowledge recall, reasoning, and code synthesis. In total, SciCode contains 338 subproblems decomposed from 80 challenging main problems. It offers optional descriptions specifying useful scientific background information and scientist-annotated gold-standard solutions and test cases for evaluation. Claude3.5-Sonnet, the best-performing model among those tested, can solve only 4.6% of the problems in the most realistic setting. We believe that SciCode demonstrates both contemporary LMs' progress towards becoming helpful scientific assistants and sheds light on the development and evaluation of scientific AI in the future.
format Preprint
id arxiv_https___arxiv_org_abs_2407_13168
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SciCode: A Research Coding Benchmark Curated by Scientists
Tian, Minyang
Gao, Luyu
Zhang, Shizhuo Dylan
Chen, Xinan
Fan, Cunwei
Guo, Xuefei
Haas, Roland
Ji, Pan
Krongchon, Kittithat
Li, Yao
Liu, Shengyan
Luo, Di
Ma, Yutao
Tong, Hao
Trinh, Kha
Tian, Chenyu
Wang, Zihan
Wu, Bohao
Xiong, Yanyu
Yin, Shengzhu
Zhu, Minhui
Lieret, Kilian
Lu, Yanxin
Liu, Genglin
Du, Yufeng
Tao, Tianhua
Press, Ofir
Callan, Jamie
Huerta, Eliu
Peng, Hao
Artificial Intelligence
Computation and Language
Since language models (LMs) now outperform average humans on many challenging tasks, it has become increasingly difficult to develop challenging, high-quality, and realistic evaluations. We address this issue by examining LMs' capabilities to generate code for solving real scientific research problems. Incorporating input from scientists and AI researchers in 16 diverse natural science sub-fields, including mathematics, physics, chemistry, biology, and materials science, we created a scientist-curated coding benchmark, SciCode. The problems in SciCode naturally factorize into multiple subproblems, each involving knowledge recall, reasoning, and code synthesis. In total, SciCode contains 338 subproblems decomposed from 80 challenging main problems. It offers optional descriptions specifying useful scientific background information and scientist-annotated gold-standard solutions and test cases for evaluation. Claude3.5-Sonnet, the best-performing model among those tested, can solve only 4.6% of the problems in the most realistic setting. We believe that SciCode demonstrates both contemporary LMs' progress towards becoming helpful scientific assistants and sheds light on the development and evaluation of scientific AI in the future.
title SciCode: A Research Coding Benchmark Curated by Scientists
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2407.13168