Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2407.13168 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911960581275648 |
|---|---|
| author | Tian, Minyang Gao, Luyu Zhang, Shizhuo Dylan Chen, Xinan Fan, Cunwei Guo, Xuefei Haas, Roland Ji, Pan Krongchon, Kittithat Li, Yao Liu, Shengyan Luo, Di Ma, Yutao Tong, Hao Trinh, Kha Tian, Chenyu Wang, Zihan Wu, Bohao Xiong, Yanyu Yin, Shengzhu Zhu, Minhui Lieret, Kilian Lu, Yanxin Liu, Genglin Du, Yufeng Tao, Tianhua Press, Ofir Callan, Jamie Huerta, Eliu Peng, Hao |
| author_facet | Tian, Minyang Gao, Luyu Zhang, Shizhuo Dylan Chen, Xinan Fan, Cunwei Guo, Xuefei Haas, Roland Ji, Pan Krongchon, Kittithat Li, Yao Liu, Shengyan Luo, Di Ma, Yutao Tong, Hao Trinh, Kha Tian, Chenyu Wang, Zihan Wu, Bohao Xiong, Yanyu Yin, Shengzhu Zhu, Minhui Lieret, Kilian Lu, Yanxin Liu, Genglin Du, Yufeng Tao, Tianhua Press, Ofir Callan, Jamie Huerta, Eliu Peng, Hao |
| contents | Since language models (LMs) now outperform average humans on many challenging tasks, it has become increasingly difficult to develop challenging, high-quality, and realistic evaluations. We address this issue by examining LMs' capabilities to generate code for solving real scientific research problems. Incorporating input from scientists and AI researchers in 16 diverse natural science sub-fields, including mathematics, physics, chemistry, biology, and materials science, we created a scientist-curated coding benchmark, SciCode. The problems in SciCode naturally factorize into multiple subproblems, each involving knowledge recall, reasoning, and code synthesis. In total, SciCode contains 338 subproblems decomposed from 80 challenging main problems. It offers optional descriptions specifying useful scientific background information and scientist-annotated gold-standard solutions and test cases for evaluation. Claude3.5-Sonnet, the best-performing model among those tested, can solve only 4.6% of the problems in the most realistic setting. We believe that SciCode demonstrates both contemporary LMs' progress towards becoming helpful scientific assistants and sheds light on the development and evaluation of scientific AI in the future. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2407_13168 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | SciCode: A Research Coding Benchmark Curated by Scientists Tian, Minyang Gao, Luyu Zhang, Shizhuo Dylan Chen, Xinan Fan, Cunwei Guo, Xuefei Haas, Roland Ji, Pan Krongchon, Kittithat Li, Yao Liu, Shengyan Luo, Di Ma, Yutao Tong, Hao Trinh, Kha Tian, Chenyu Wang, Zihan Wu, Bohao Xiong, Yanyu Yin, Shengzhu Zhu, Minhui Lieret, Kilian Lu, Yanxin Liu, Genglin Du, Yufeng Tao, Tianhua Press, Ofir Callan, Jamie Huerta, Eliu Peng, Hao Artificial Intelligence Computation and Language Since language models (LMs) now outperform average humans on many challenging tasks, it has become increasingly difficult to develop challenging, high-quality, and realistic evaluations. We address this issue by examining LMs' capabilities to generate code for solving real scientific research problems. Incorporating input from scientists and AI researchers in 16 diverse natural science sub-fields, including mathematics, physics, chemistry, biology, and materials science, we created a scientist-curated coding benchmark, SciCode. The problems in SciCode naturally factorize into multiple subproblems, each involving knowledge recall, reasoning, and code synthesis. In total, SciCode contains 338 subproblems decomposed from 80 challenging main problems. It offers optional descriptions specifying useful scientific background information and scientist-annotated gold-standard solutions and test cases for evaluation. Claude3.5-Sonnet, the best-performing model among those tested, can solve only 4.6% of the problems in the most realistic setting. We believe that SciCode demonstrates both contemporary LMs' progress towards becoming helpful scientific assistants and sheds light on the development and evaluation of scientific AI in the future. |
| title | SciCode: A Research Coding Benchmark Curated by Scientists |
| topic | Artificial Intelligence Computation and Language |
| url | https://arxiv.org/abs/2407.13168 |