MedAgentGym: A Scalable Agentic Training Environment for Code-Centric Reasoning in Biomedical Data Science

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Ran, Zhuang, Yuchen, Zhong, Yishan, Yu, Yue, Wang, Zifeng, Tang, Xiangru, Wu, Hang, Wang, May D., Ruan, Peifeng, Yang, Donghan, Wang, Tao, Xiao, Guanghua, Liu, Xin, Yang, Carl, Xie, Yang, Shi, Wenqi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912628592345088
author Xu, Ran
Zhuang, Yuchen
Zhong, Yishan
Yu, Yue
Wang, Zifeng
Tang, Xiangru
Wu, Hang
Wang, May D.
Ruan, Peifeng
Yang, Donghan
Wang, Tao
Xiao, Guanghua
Liu, Xin
Yang, Carl
Xie, Yang
Shi, Wenqi
author_facet Xu, Ran
Zhuang, Yuchen
Zhong, Yishan
Yu, Yue
Wang, Zifeng
Tang, Xiangru
Wu, Hang
Wang, May D.
Ruan, Peifeng
Yang, Donghan
Wang, Tao
Xiao, Guanghua
Liu, Xin
Yang, Carl
Xie, Yang
Shi, Wenqi
contents We introduce MedAgentGym, a scalable and interactive training environment designed to enhance coding-based biomedical reasoning capabilities in large language model (LLM) agents. MedAgentGym comprises 72,413 task instances across 129 categories derived from 12 authentic real-world biomedical scenarios. Tasks are encapsulated within executable sandbox environments, each featuring detailed task specifications, interactive feedback mechanisms, verifiable ground truth annotations, and scalable training trajectory generation. Extensive benchmarking of 29 LLMs reveals substantial performance disparities in biomedical data science between commercial and open-source LLMs. Leveraging efficient multi-threaded and multi-turn trajectory sampling in MedAgentGym, Med-Copilot achieves performance gains of +43.02% and +45.28% from offline and online reinforcement learning, respectively, demonstrating MedAgentGym as an effective training ground while establishing itself as a cost-effective, privacy-preserving alternative competitive with proprietary LLMs (gpt-4o). By offering a unified execution environment with a comprehensive benchmark and accessible, extensible training resources, MedAgentGym delivers an integrated platform to develop LLM-based coding assistants for advanced biomedical data science.
format Preprint
id arxiv_https___arxiv_org_abs_2506_04405
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MedAgentGym: A Scalable Agentic Training Environment for Code-Centric Reasoning in Biomedical Data Science
Xu, Ran
Zhuang, Yuchen
Zhong, Yishan
Yu, Yue
Wang, Zifeng
Tang, Xiangru
Wu, Hang
Wang, May D.
Ruan, Peifeng
Yang, Donghan
Wang, Tao
Xiao, Guanghua
Liu, Xin
Yang, Carl
Xie, Yang
Shi, Wenqi
Computation and Language
Artificial Intelligence
Machine Learning
We introduce MedAgentGym, a scalable and interactive training environment designed to enhance coding-based biomedical reasoning capabilities in large language model (LLM) agents. MedAgentGym comprises 72,413 task instances across 129 categories derived from 12 authentic real-world biomedical scenarios. Tasks are encapsulated within executable sandbox environments, each featuring detailed task specifications, interactive feedback mechanisms, verifiable ground truth annotations, and scalable training trajectory generation. Extensive benchmarking of 29 LLMs reveals substantial performance disparities in biomedical data science between commercial and open-source LLMs. Leveraging efficient multi-threaded and multi-turn trajectory sampling in MedAgentGym, Med-Copilot achieves performance gains of +43.02% and +45.28% from offline and online reinforcement learning, respectively, demonstrating MedAgentGym as an effective training ground while establishing itself as a cost-effective, privacy-preserving alternative competitive with proprietary LLMs (gpt-4o). By offering a unified execution environment with a comprehensive benchmark and accessible, extensible training resources, MedAgentGym delivers an integrated platform to develop LLM-based coding assistants for advanced biomedical data science.
title MedAgentGym: A Scalable Agentic Training Environment for Code-Centric Reasoning in Biomedical Data Science
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2506.04405