JunoBench: A Benchmark Dataset of Crashes in Python Machine Learning Jupyter Notebooks

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Yiran, López, José Antonio Hernández, Nilsson, Ulf, Varró, Dániel
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914518989275136
author Wang, Yiran
López, José Antonio Hernández
Nilsson, Ulf
Varró, Dániel
author_facet Wang, Yiran
López, José Antonio Hernández
Nilsson, Ulf
Varró, Dániel
contents Jupyter notebooks are widely used for machine learning (ML) prototyping. Yet, few debugging tools are designed for ML code in notebooks, partly, due to the lack of benchmarks. We introduce JunoBench, the first benchmark dataset of real-world crashes in Python-based ML notebooks. JunoBench includes 111 curated and reproducible crashes with verified fixes from public Kaggle notebooks, covering popular ML libraries (e.g., TensorFlow/Keras, PyTorch, Scikit-learn) and notebook-specific out-of-order execution errors. JunoBench ensures reproducibility and ease of use through a unified environment that reliably reproduces all crashes. By providing realistic crashes, their resolutions, richly annotated labels of crash characteristics, and natural-language diagnostic annotations, JunoBench facilitates research on bug detection, localization, diagnosis, and repair in notebook-based ML development.
format Preprint
id arxiv_https___arxiv_org_abs_2510_18013
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle JunoBench: A Benchmark Dataset of Crashes in Python Machine Learning Jupyter Notebooks
Wang, Yiran
López, José Antonio Hernández
Nilsson, Ulf
Varró, Dániel
Software Engineering
Jupyter notebooks are widely used for machine learning (ML) prototyping. Yet, few debugging tools are designed for ML code in notebooks, partly, due to the lack of benchmarks. We introduce JunoBench, the first benchmark dataset of real-world crashes in Python-based ML notebooks. JunoBench includes 111 curated and reproducible crashes with verified fixes from public Kaggle notebooks, covering popular ML libraries (e.g., TensorFlow/Keras, PyTorch, Scikit-learn) and notebook-specific out-of-order execution errors. JunoBench ensures reproducibility and ease of use through a unified environment that reliably reproduces all crashes. By providing realistic crashes, their resolutions, richly annotated labels of crash characteristics, and natural-language diagnostic annotations, JunoBench facilitates research on bug detection, localization, diagnosis, and repair in notebook-based ML development.
title JunoBench: A Benchmark Dataset of Crashes in Python Machine Learning Jupyter Notebooks
topic Software Engineering
url https://arxiv.org/abs/2510.18013