ML2B: Multi-Lingual ML Benchmark For AutoML

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Trofimova, Ekaterina, Shamina, Zosia, Selifanova, Maria, Zaitsev, Artem, Savchuk, Remi, Minets, Maxim, Ozerova, Daria, Sataev, Emil, Zuenko, Denis, Ustyuzhanin, Andrey E.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914077229449216
author Trofimova, Ekaterina
Shamina, Zosia
Selifanova, Maria
Zaitsev, Artem
Savchuk, Remi
Minets, Maxim
Ozerova, Daria
Sataev, Emil
Zuenko, Denis
Ustyuzhanin, Andrey E.
author_facet Trofimova, Ekaterina
Shamina, Zosia
Selifanova, Maria
Zaitsev, Artem
Savchuk, Remi
Minets, Maxim
Ozerova, Daria
Sataev, Emil
Zuenko, Denis
Ustyuzhanin, Andrey E.
contents Large language models (LLMs) have recently demonstrated strong capabilities in generating machine learning (ML) code, enabling end-to-end pipeline construction from natural language instructions. However, existing benchmarks for ML code generation are mainly restricted to English, overlooking the global and multilingual nature of ML research and practice. To address this gap, we present ML2B, the first benchmark for evaluating multilingual ML code generation. ML2B consists of 30 Kaggle competitions translated into 13 natural languages, covering tabular, text, and image data types, with structured metadata and validated human-reviewed translations. For evaluation, we employ AIDE, an automated framework for end-to-end assessment of data science pipelines, and provide insights into cross-lingual model performance. Our results reveal substantial 15-45% performance degradation on non-English tasks, highlighting critical challenges in multilingual representation learning for code generation. The benchmark, evaluation framework, and comprehensive results are made available through our GitHub repository to facilitate future research in multilingual ML code generation: https://github.com/enaix/ml2b.
format Preprint
id arxiv_https___arxiv_org_abs_2509_22768
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ML2B: Multi-Lingual ML Benchmark For AutoML
Trofimova, Ekaterina
Shamina, Zosia
Selifanova, Maria
Zaitsev, Artem
Savchuk, Remi
Minets, Maxim
Ozerova, Daria
Sataev, Emil
Zuenko, Denis
Ustyuzhanin, Andrey E.
Computation and Language
Large language models (LLMs) have recently demonstrated strong capabilities in generating machine learning (ML) code, enabling end-to-end pipeline construction from natural language instructions. However, existing benchmarks for ML code generation are mainly restricted to English, overlooking the global and multilingual nature of ML research and practice. To address this gap, we present ML2B, the first benchmark for evaluating multilingual ML code generation. ML2B consists of 30 Kaggle competitions translated into 13 natural languages, covering tabular, text, and image data types, with structured metadata and validated human-reviewed translations. For evaluation, we employ AIDE, an automated framework for end-to-end assessment of data science pipelines, and provide insights into cross-lingual model performance. Our results reveal substantial 15-45% performance degradation on non-English tasks, highlighting critical challenges in multilingual representation learning for code generation. The benchmark, evaluation framework, and comprehensive results are made available through our GitHub repository to facilitate future research in multilingual ML code generation: https://github.com/enaix/ml2b.
title ML2B: Multi-Lingual ML Benchmark For AutoML
topic Computation and Language
url https://arxiv.org/abs/2509.22768