Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Li, Loka, Zhang, Duzhen, Du, Xingbo, Song, Leonard, Wang, Zixiao, Aukenov, Assanali, Thomas, Noel, Sailaukan, Shakhnazar, Yang, Yonghan, Chen, Feilong, Dong, Jiahua, Zhang, Kun, Zhang, Bin, Song, Le
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:https://arxiv.org/abs/2605.15766
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911688196882432
author Li, Loka
Zhang, Duzhen
Du, Xingbo
Song, Leonard
Wang, Zixiao
Aukenov, Assanali
Thomas, Noel
Sailaukan, Shakhnazar
Yang, Yonghan
Chen, Feilong
Dong, Jiahua
Zhang, Kun
Zhang, Bin
Song, Le
author_facet Li, Loka
Zhang, Duzhen
Du, Xingbo
Song, Leonard
Wang, Zixiao
Aukenov, Assanali
Thomas, Noel
Sailaukan, Shakhnazar
Yang, Yonghan
Chen, Feilong
Dong, Jiahua
Zhang, Kun
Zhang, Bin
Song, Le
contents Large language model (LLM) agents are increasingly capable of automating components of machine learning development, yet existing biomedical benchmarks mainly focus on question answering, reasoning, and tool usage, or evaluate only narrow aspects of biomedical ML coding. We present BioXArena, a biomedical machine learning benchmark designed to evaluate whether agents can generate task-specific model training pipelines for heterogeneous and multi-modal biomedical datasets. BioXArena contains 76 end-to-end tasks across 9 domains, including sequence modeling, single-cell analysis, structural biology, network biology, chemical biology, perturbation dynamics, phenotype-disease modeling, biomedical imaging, and text-integrated learning. Each task is curated from primary biomedical sources into a unified evaluation framework with hidden labels, held-out graders, and biology-aware metrics normalized to a 0 to 1 scale. Agents are required to write executable code, train predictive models, and generate submissions for private test samples. Most tasks involve multiple input modalities, including tabular data, images, natural language, molecular sequences, omics matrices, and protein structures. We evaluate 11 agent configurations in a standardized 2-hour single-GPU environment. MLEvolve with Gemini-3.1-Pro achieves the highest average score of 0.666, followed by GPT-5.4 with 0.636, while no single agent consistently dominates across all domains. We additionally perform extensive ablation studies, robustness evaluations, scaling analyses, cost analyses, and failure-mode investigations to better understand how model backbones, agent scaffolds, inference budgets, and biomedical domains influence BioML coding performance. We will publicly release all benchmark tasks, graders, execution runners, leaderboard results, and agent trajectories.
format Preprint
id arxiv_https___arxiv_org_abs_2605_15766
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle BioXArena: Benchmarking LLM Agents on Multi-Modal Biomedical Machine Learning Tasks
Li, Loka
Zhang, Duzhen
Du, Xingbo
Song, Leonard
Wang, Zixiao
Aukenov, Assanali
Thomas, Noel
Sailaukan, Shakhnazar
Yang, Yonghan
Chen, Feilong
Dong, Jiahua
Zhang, Kun
Zhang, Bin
Song, Le
Computational Engineering, Finance, and Science
Large language model (LLM) agents are increasingly capable of automating components of machine learning development, yet existing biomedical benchmarks mainly focus on question answering, reasoning, and tool usage, or evaluate only narrow aspects of biomedical ML coding. We present BioXArena, a biomedical machine learning benchmark designed to evaluate whether agents can generate task-specific model training pipelines for heterogeneous and multi-modal biomedical datasets. BioXArena contains 76 end-to-end tasks across 9 domains, including sequence modeling, single-cell analysis, structural biology, network biology, chemical biology, perturbation dynamics, phenotype-disease modeling, biomedical imaging, and text-integrated learning. Each task is curated from primary biomedical sources into a unified evaluation framework with hidden labels, held-out graders, and biology-aware metrics normalized to a 0 to 1 scale. Agents are required to write executable code, train predictive models, and generate submissions for private test samples. Most tasks involve multiple input modalities, including tabular data, images, natural language, molecular sequences, omics matrices, and protein structures. We evaluate 11 agent configurations in a standardized 2-hour single-GPU environment. MLEvolve with Gemini-3.1-Pro achieves the highest average score of 0.666, followed by GPT-5.4 with 0.636, while no single agent consistently dominates across all domains. We additionally perform extensive ablation studies, robustness evaluations, scaling analyses, cost analyses, and failure-mode investigations to better understand how model backbones, agent scaffolds, inference budgets, and biomedical domains influence BioML coding performance. We will publicly release all benchmark tasks, graders, execution runners, leaderboard results, and agent trajectories.
title BioXArena: Benchmarking LLM Agents on Multi-Modal Biomedical Machine Learning Tasks
topic Computational Engineering, Finance, and Science
url https://arxiv.org/abs/2605.15766