MLZero: A Multi-Agent System for End-to-end Machine Learning Automation

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Fang, Haoyang, Han, Boran, Erickson, Nick, Zhang, Xiyuan, Zhou, Su, Dagar, Anirudh, Zhang, Jiani, Turkmen, Ali Caner, Hu, Cuixiong, Rangwala, Huzefa, Wu, Ying Nian, Wang, Bernie, Karypis, George
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915293669883904
author Fang, Haoyang
Han, Boran
Erickson, Nick
Zhang, Xiyuan
Zhou, Su
Dagar, Anirudh
Zhang, Jiani
Turkmen, Ali Caner
Hu, Cuixiong
Rangwala, Huzefa
Wu, Ying Nian
Wang, Bernie
Karypis, George
author_facet Fang, Haoyang
Han, Boran
Erickson, Nick
Zhang, Xiyuan
Zhou, Su
Dagar, Anirudh
Zhang, Jiani
Turkmen, Ali Caner
Hu, Cuixiong
Rangwala, Huzefa
Wu, Ying Nian
Wang, Bernie
Karypis, George
contents Existing AutoML systems have advanced the automation of machine learning (ML); however, they still require substantial manual configuration and expert input, particularly when handling multimodal data. We introduce MLZero, a novel multi-agent framework powered by Large Language Models (LLMs) that enables end-to-end ML automation across diverse data modalities with minimal human intervention. A cognitive perception module is first employed, transforming raw multimodal inputs into perceptual context that effectively guides the subsequent workflow. To address key limitations of LLMs, such as hallucinated code generation and outdated API knowledge, we enhance the iterative code generation process with semantic and episodic memory. MLZero demonstrates superior performance on MLE-Bench Lite, outperforming all competitors in both success rate and solution quality, securing six gold medals. Additionally, when evaluated on our Multimodal AutoML Agent Benchmark, which includes 25 more challenging tasks spanning diverse data modalities, MLZero outperforms the competing methods by a large margin with a success rate of 0.92 (+263.6\%) and an average rank of 2.28. Our approach maintains its robust effectiveness even with a compact 8B LLM, outperforming full-size systems from existing solutions.
format Preprint
id arxiv_https___arxiv_org_abs_2505_13941
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MLZero: A Multi-Agent System for End-to-end Machine Learning Automation
Fang, Haoyang
Han, Boran
Erickson, Nick
Zhang, Xiyuan
Zhou, Su
Dagar, Anirudh
Zhang, Jiani
Turkmen, Ali Caner
Hu, Cuixiong
Rangwala, Huzefa
Wu, Ying Nian
Wang, Bernie
Karypis, George
Multiagent Systems
Artificial Intelligence
Computation and Language
Machine Learning
Existing AutoML systems have advanced the automation of machine learning (ML); however, they still require substantial manual configuration and expert input, particularly when handling multimodal data. We introduce MLZero, a novel multi-agent framework powered by Large Language Models (LLMs) that enables end-to-end ML automation across diverse data modalities with minimal human intervention. A cognitive perception module is first employed, transforming raw multimodal inputs into perceptual context that effectively guides the subsequent workflow. To address key limitations of LLMs, such as hallucinated code generation and outdated API knowledge, we enhance the iterative code generation process with semantic and episodic memory. MLZero demonstrates superior performance on MLE-Bench Lite, outperforming all competitors in both success rate and solution quality, securing six gold medals. Additionally, when evaluated on our Multimodal AutoML Agent Benchmark, which includes 25 more challenging tasks spanning diverse data modalities, MLZero outperforms the competing methods by a large margin with a success rate of 0.92 (+263.6\%) and an average rank of 2.28. Our approach maintains its robust effectiveness even with a compact 8B LLM, outperforming full-size systems from existing solutions.
title MLZero: A Multi-Agent System for End-to-end Machine Learning Automation
topic Multiagent Systems
Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2505.13941