Saved in:
Bibliographic Details
Main Authors: Lyu, Yougang, Zhang, Xi, Yi, Xinhao, Zhao, Yuyue, Guo, Shuyu, Hu, Wenxiang, Piotrowski, Jan, Kaliski, Jakub, Urbani, Jacopo, Meng, Zaiqiao, Zhou, Lun, Yan, Xiaohui
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2603.08127
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917324392497152
author Lyu, Yougang
Zhang, Xi
Yi, Xinhao
Zhao, Yuyue
Guo, Shuyu
Hu, Wenxiang
Piotrowski, Jan
Kaliski, Jakub
Urbani, Jacopo
Meng, Zaiqiao
Zhou, Lun
Yan, Xiaohui
author_facet Lyu, Yougang
Zhang, Xi
Yi, Xinhao
Zhao, Yuyue
Guo, Shuyu
Hu, Wenxiang
Piotrowski, Jan
Kaliski, Jakub
Urbani, Jacopo
Meng, Zaiqiao
Zhou, Lun
Yan, Xiaohui
contents The increasing adoption of Large Language Models (LLMs) has enabled AI scientists to perform complex end-to-end scientific discovery tasks requiring coordination of specialized roles, including idea generation and experimental execution. However, most state-of-the-art AI scientist systems rely on static, hand-designed pipelines and fail to adapt based on accumulated interaction histories. As a result, these systems overlook promising research directions, repeat failed experiments, and pursue infeasible ideas. To address this, we introduce EvoScientist, an evolving multi-agent AI scientist framework that continuously improves research strategies through persistent memory and self-evolution. EvoScientist comprises three specialized agents: a Researcher Agent (RA) for scientific idea generation, an Engineer Agent (EA) for experiment implementation and execution, and an Evolution Manager Agent (EMA) that distills insights from prior interactions into reusable knowledge. EvoScientist contains two persistent memory modules: (i) an ideation memory, which summarizes feasible research directions from top-ranked ideas while recording previously unsuccessful directions; and (ii) an experimentation memory, which captures effective data processing and model training strategies derived from code search trajectories and best-performing implementations. These modules enable the RA and EA to retrieve relevant prior strategies, improving idea quality and code execution success rates over time. Experiments show that EvoScientist outperforms 7 open-source and commercial state-of-the-art systems in scientific idea generation, achieving higher novelty, feasibility, relevance, and clarity via automatic and human evaluation. EvoScientist also substantially improves code execution success rates through multi-agent evolution, demonstrating persistent memory's effectiveness for end-to-end scientific discovery.
format Preprint
id arxiv_https___arxiv_org_abs_2603_08127
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle EvoScientist: Towards Multi-Agent Evolving AI Scientists for End-to-End Scientific Discovery
Lyu, Yougang
Zhang, Xi
Yi, Xinhao
Zhao, Yuyue
Guo, Shuyu
Hu, Wenxiang
Piotrowski, Jan
Kaliski, Jakub
Urbani, Jacopo
Meng, Zaiqiao
Zhou, Lun
Yan, Xiaohui
Computation and Language
The increasing adoption of Large Language Models (LLMs) has enabled AI scientists to perform complex end-to-end scientific discovery tasks requiring coordination of specialized roles, including idea generation and experimental execution. However, most state-of-the-art AI scientist systems rely on static, hand-designed pipelines and fail to adapt based on accumulated interaction histories. As a result, these systems overlook promising research directions, repeat failed experiments, and pursue infeasible ideas. To address this, we introduce EvoScientist, an evolving multi-agent AI scientist framework that continuously improves research strategies through persistent memory and self-evolution. EvoScientist comprises three specialized agents: a Researcher Agent (RA) for scientific idea generation, an Engineer Agent (EA) for experiment implementation and execution, and an Evolution Manager Agent (EMA) that distills insights from prior interactions into reusable knowledge. EvoScientist contains two persistent memory modules: (i) an ideation memory, which summarizes feasible research directions from top-ranked ideas while recording previously unsuccessful directions; and (ii) an experimentation memory, which captures effective data processing and model training strategies derived from code search trajectories and best-performing implementations. These modules enable the RA and EA to retrieve relevant prior strategies, improving idea quality and code execution success rates over time. Experiments show that EvoScientist outperforms 7 open-source and commercial state-of-the-art systems in scientific idea generation, achieving higher novelty, feasibility, relevance, and clarity via automatic and human evaluation. EvoScientist also substantially improves code execution success rates through multi-agent evolution, demonstrating persistent memory's effectiveness for end-to-end scientific discovery.
title EvoScientist: Towards Multi-Agent Evolving AI Scientists for End-to-End Scientific Discovery
topic Computation and Language
url https://arxiv.org/abs/2603.08127