JunoBench: A Benchmark Dataset of Crashes in Python Machine Learning Jupyter Notebooks
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Yiran, López, José Antonio Hernández, Nilsson, Ulf, Varró, Dániel |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Why do Machine Learning Notebooks Crash? An Empirical Study on Public Python Jupyter Notebooks
by: Wang, Yiran, et al.
Published: (2024)
by: Wang, Yiran, et al.
Published: (2024)
Runtime-Augmented LLMs for Crash Detection and Diagnosis in ML Notebooks
by: Wang, Yiran, et al.
Published: (2026)
by: Wang, Yiran, et al.
Published: (2026)
Understanding Feedback Mechanisms in Machine Learning Jupyter Notebooks
by: Shome, Arumoy, et al.
Published: (2024)
by: Shome, Arumoy, et al.
Published: (2024)
Method Names in Jupyter Notebooks: An Exploratory Study
by: Wong, Carol, et al.
Published: (2025)
by: Wong, Carol, et al.
Published: (2025)
Mining the Characteristics of Jupyter Notebooks in Data Science Projects
by: Choetkiertikul, Morakot, et al.
Published: (2023)
by: Choetkiertikul, Morakot, et al.
Published: (2023)
Typhon: Automatic Recommendation of Relevant Code Cells in Jupyter Notebooks
by: Ragkhitwetsagul, Chaiyong, et al.
Published: (2024)
by: Ragkhitwetsagul, Chaiyong, et al.
Published: (2024)
Similarity-Based Assessment of Computational Reproducibility in Jupyter Notebooks
by: Hossain, A S M Shahadat, et al.
Published: (2025)
by: Hossain, A S M Shahadat, et al.
Published: (2025)
Analysing Python Machine Learning Notebooks with Moose
by: Mignard, Marius, et al.
Published: (2025)
by: Mignard, Marius, et al.
Published: (2025)
Observing Fine-Grained Changes in Jupyter Notebooks During Development Time
by: Titov, Sergey, et al.
Published: (2025)
by: Titov, Sergey, et al.
Published: (2025)
Improving Quantum Developer Experience with Kubernetes and Jupyter Notebooks
by: Kinanen, Otso, et al.
Published: (2024)
by: Kinanen, Otso, et al.
Published: (2024)
Containing the Reproducibility Gap: Automated Repository-Level Containerization for Scholarly Jupyter Notebooks
by: Samuel, Sheeba, et al.
Published: (2026)
by: Samuel, Sheeba, et al.
Published: (2026)
A Flexible Cell Classification for ML Projects in Jupyter Notebooks
by: Perez, Miguel, et al.
Published: (2024)
by: Perez, Miguel, et al.
Published: (2024)
Generative AI in Simulation-Based Test Environments for Large-Scale Cyber-Physical Systems: An Industrial Study
by: Sadrnezhaad, Masoud, et al.
Published: (2025)
by: Sadrnezhaad, Masoud, et al.
Published: (2025)
A Systematic Literature Review of Software Engineering Research on Jupyter Notebook
by: Siddik, Md Saeed, et al.
Published: (2025)
by: Siddik, Md Saeed, et al.
Published: (2025)
How Scientists Use Jupyter Notebooks: Goals, Quality Attributes, and Opportunities
by: Huang, Ruanqianqian, et al.
Published: (2025)
by: Huang, Ruanqianqian, et al.
Published: (2025)
Hierarchical Evaluation of Software Design Capabilities of Large Language Models of Code
by: Saad, Mootez, et al.
Published: (2025)
by: Saad, Mootez, et al.
Published: (2025)
ALPINE: An adaptive language-agnostic pruning method for language models for code
by: Saad, Mootez, et al.
Published: (2024)
by: Saad, Mootez, et al.
Published: (2024)
On Inter-dataset Code Duplication and Data Leakage in Large Language Models
by: López, José Antonio Hernández, et al.
Published: (2024)
by: López, José Antonio Hernández, et al.
Published: (2024)
DyPyBench: A Benchmark of Executable Python Software
by: Bouzenia, Islem, et al.
Published: (2024)
by: Bouzenia, Islem, et al.
Published: (2024)
CrashJS: A NodeJS Benchmark for Automated Crash Reproduction
by: Oliver, Philip, et al.
Published: (2024)
by: Oliver, Philip, et al.
Published: (2024)
Static Analysis Driven Enhancements for Comprehension in Machine Learning Notebooks
by: Venkatesh, Ashwin Prasad Shivarpatna, et al.
Published: (2023)
by: Venkatesh, Ashwin Prasad Shivarpatna, et al.
Published: (2023)
SENAI: Towards Software Engineering Native Generative Artificial Intelligence
by: Saad, Mootez, et al.
Published: (2025)
by: Saad, Mootez, et al.
Published: (2025)
KGym: A Platform and Dataset to Benchmark Large Language Models on Linux Kernel Crash Resolution
by: Mathai, Alex, et al.
Published: (2024)
by: Mathai, Alex, et al.
Published: (2024)
LeakageDetector 2.0: Analyzing Data Leakage in Jupyter-Driven Machine Learning Pipelines
by: Truong, Owen, et al.
Published: (2025)
by: Truong, Owen, et al.
Published: (2025)
Automated Modernization of Machine Learning Engineering Notebooks for Reproducibility
by: Jin, Bihui, et al.
Published: (2026)
by: Jin, Bihui, et al.
Published: (2026)
The Power of Types: Exploring the Impact of Type Checking on Neural Bug Detection in Dynamically Typed Languages
by: Chen, Boqi, et al.
Published: (2024)
by: Chen, Boqi, et al.
Published: (2024)
A Regression Testing Framework with Automated Assertion Generation for Machine Learning Notebooks
by: Yao, Yingao Elaine, et al.
Published: (2025)
by: Yao, Yingao Elaine, et al.
Published: (2025)
Characterising Bugs in Jupyter Platform
by: Tang, Yutian, et al.
Published: (2025)
by: Tang, Yutian, et al.
Published: (2025)
SHERPA: A Model-Driven Framework for Large Language Model Execution
by: Chen, Boqi, et al.
Published: (2025)
by: Chen, Boqi, et al.
Published: (2025)
Themisto: Jupyter-Based Runtime Benchmark
by: Grotov, Konstantin, et al.
Published: (2025)
by: Grotov, Konstantin, et al.
Published: (2025)
Concretization of Abstract Traffic Scene Specifications Using Metaheuristic Search
by: Babikian, Aren A., et al.
Published: (2023)
by: Babikian, Aren A., et al.
Published: (2023)
CodeFuse-CR-Bench: A Comprehensiveness-aware Benchmark for End-to-End Code Review Evaluation in Python Projects
by: Guo, Hanyang, et al.
Published: (2025)
by: Guo, Hanyang, et al.
Published: (2025)
Exploring the Jupyter Ecosystem: An Empirical Study of Bugs and Vulnerabilities
by: Jiang, Wenyuan, et al.
Published: (2025)
by: Jiang, Wenyuan, et al.
Published: (2025)
Outrunning LLM Cutoffs: A Live Kernel Crash Resolution Benchmark for All
by: Huang, Chenxi, et al.
Published: (2026)
by: Huang, Chenxi, et al.
Published: (2026)
Detecting Refactoring Commits in Machine Learning Python Projects: A Machine Learning-Based Approach
by: Noei, Shayan, et al.
Published: (2024)
by: Noei, Shayan, et al.
Published: (2024)
Finding the Needle in the Crash Stack: Industrial-Scale Crash Root Cause Localization with AutoCrashFL
by: Kang, Sungmin, et al.
Published: (2025)
by: Kang, Sungmin, et al.
Published: (2025)
A Study of Scientific Computational Notebook Quality
by: Kashiwa, Shun, et al.
Published: (2026)
by: Kashiwa, Shun, et al.
Published: (2026)
Deep-Bench: Deep Learning Benchmark Dataset for Code Generation
by: Daghighfarsoodeh, Alireza, et al.
Published: (2025)
by: Daghighfarsoodeh, Alireza, et al.
Published: (2025)
RealBench: A Repo-Level Code Generation Benchmark Aligned with Real-World Software Development Practices
by: Li, Jia, et al.
Published: (2026)
by: Li, Jia, et al.
Published: (2026)
LLM-based Satisfiability Checking of String Requirements by Consistent Data and Checker Generation
by: Chen, Boqi, et al.
Published: (2025)
by: Chen, Boqi, et al.
Published: (2025)
Similar Items
-
Why do Machine Learning Notebooks Crash? An Empirical Study on Public Python Jupyter Notebooks
by: Wang, Yiran, et al.
Published: (2024) -
Runtime-Augmented LLMs for Crash Detection and Diagnosis in ML Notebooks
by: Wang, Yiran, et al.
Published: (2026) -
Understanding Feedback Mechanisms in Machine Learning Jupyter Notebooks
by: Shome, Arumoy, et al.
Published: (2024) -
Method Names in Jupyter Notebooks: An Exploratory Study
by: Wong, Carol, et al.
Published: (2025) -
Mining the Characteristics of Jupyter Notebooks in Data Science Projects
by: Choetkiertikul, Morakot, et al.
Published: (2023)