Measuring Scientific Capabilities of Language Models with a Systems Biology Dry Lab

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Duan, Haonan, Lu, Stephen Zhewen, Harrigan, Caitlin Fiona, Desai, Nishkrit, Lu, Jiarui, Koziarski, Michał, Cotta, Leonardo, Maddison, Chris J.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918090915184640
author Duan, Haonan
Lu, Stephen Zhewen
Harrigan, Caitlin Fiona
Desai, Nishkrit
Lu, Jiarui
Koziarski, Michał
Cotta, Leonardo
Maddison, Chris J.
author_facet Duan, Haonan
Lu, Stephen Zhewen
Harrigan, Caitlin Fiona
Desai, Nishkrit
Lu, Jiarui
Koziarski, Michał
Cotta, Leonardo
Maddison, Chris J.
contents Designing experiments and result interpretations are core scientific competencies, particularly in biology, where researchers perturb complex systems to uncover the underlying systems. Recent efforts to evaluate the scientific capabilities of large language models (LLMs) fail to test these competencies because wet-lab experimentation is prohibitively expensive: in expertise, time and equipment. We introduce SciGym, a first-in-class benchmark that assesses LLMs' iterative experiment design and analysis abilities in open-ended scientific discovery tasks. SciGym overcomes the challenge of wet-lab costs by running a dry lab of biological systems. These models, encoded in Systems Biology Markup Language, are efficient for generating simulated data, making them ideal testbeds for experimentation on realistically complex systems. We evaluated six frontier LLMs on 137 small systems, and released a total of 350 systems. Our evaluation shows that while more capable models demonstrated superior performance, all models' performance declined significantly as system complexity increased, suggesting substantial room for improvement in the scientific capabilities of LLM agents.
format Preprint
id arxiv_https___arxiv_org_abs_2507_02083
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Measuring Scientific Capabilities of Language Models with a Systems Biology Dry Lab
Duan, Haonan
Lu, Stephen Zhewen
Harrigan, Caitlin Fiona
Desai, Nishkrit
Lu, Jiarui
Koziarski, Michał
Cotta, Leonardo
Maddison, Chris J.
Artificial Intelligence
Designing experiments and result interpretations are core scientific competencies, particularly in biology, where researchers perturb complex systems to uncover the underlying systems. Recent efforts to evaluate the scientific capabilities of large language models (LLMs) fail to test these competencies because wet-lab experimentation is prohibitively expensive: in expertise, time and equipment. We introduce SciGym, a first-in-class benchmark that assesses LLMs' iterative experiment design and analysis abilities in open-ended scientific discovery tasks. SciGym overcomes the challenge of wet-lab costs by running a dry lab of biological systems. These models, encoded in Systems Biology Markup Language, are efficient for generating simulated data, making them ideal testbeds for experimentation on realistically complex systems. We evaluated six frontier LLMs on 137 small systems, and released a total of 350 systems. Our evaluation shows that while more capable models demonstrated superior performance, all models' performance declined significantly as system complexity increased, suggesting substantial room for improvement in the scientific capabilities of LLM agents.
title Measuring Scientific Capabilities of Language Models with a Systems Biology Dry Lab
topic Artificial Intelligence
url https://arxiv.org/abs/2507.02083