MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility
Fuente:
arXiv
Saved in:
| Main Authors: | Gaddipati, Sasi Kiran, Muhammed, Diyana, Keya, Farhana, Rabby, Gollam, Auer, Sören |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AISSISTANT: Human-AI Collaborative Review and Perspective Research Workflows in Data Science
by: Gaddipati, Sasi Kiran, et al.
Published: (2025)
by: Gaddipati, Sasi Kiran, et al.
Published: (2025)
MC-NEST: Enhancing Mathematical Reasoning in Large Language Models leveraging a Monte Carlo Self-Refine Tree
by: Rabby, Gollam, et al.
Published: (2024)
by: Rabby, Gollam, et al.
Published: (2024)
SelfCheck-Eval: A Multi-Module Framework for Zero-Resource Hallucination Detection in Large Language Models
by: Muhammed, Diyana, et al.
Published: (2025)
by: Muhammed, Diyana, et al.
Published: (2025)
Iterative Hypothesis Generation for Scientific Discovery with Monte Carlo Nash Equilibrium Self-Refining Trees
by: Rabby, Gollam, et al.
Published: (2025)
by: Rabby, Gollam, et al.
Published: (2025)
Fine-tuning and Prompt Engineering with Cognitive Knowledge Graphs for Scholarly Knowledge Organization
by: Rabby, Gollam, et al.
Published: (2024)
by: Rabby, Gollam, et al.
Published: (2024)
SCI-IDEA: Context-Aware Scientific Ideation Using Token and Sentence Embeddings
by: Keya, Farhana, et al.
Published: (2025)
by: Keya, Farhana, et al.
Published: (2025)
Towards AI-Supported Research: a Vision of the TIB AIssistant
by: Auer, Sören, et al.
Published: (2025)
by: Auer, Sören, et al.
Published: (2025)
SLEEPING-DISCO 9M: A large-scale pre-training dataset for generative music modeling
by: Ahmed, Tawsif, et al.
Published: (2025)
by: Ahmed, Tawsif, et al.
Published: (2025)
NeuroSym-BioCAT: Leveraging Neuro-Symbolic Methods for Biomedical Scholarly Document Categorization and Question Answering
by: Zamil, Parvez, et al.
Published: (2024)
by: Zamil, Parvez, et al.
Published: (2024)
Project Alexandria: Towards Freeing Scientific Knowledge from Copyright Burdens via LLMs
by: Schuhmann, Christoph, et al.
Published: (2025)
by: Schuhmann, Christoph, et al.
Published: (2025)
What Do Machine Learning Researchers Mean by "Reproducible"?
by: Raff, Edward, et al.
Published: (2024)
by: Raff, Edward, et al.
Published: (2024)
Predicting Diabetes Using Machine Learning: A Comparative Study of Classifiers
by: Hasan, Mahade, et al.
Published: (2025)
by: Hasan, Mahade, et al.
Published: (2025)
EmoNet-Voice: A Fine-Grained, Expert-Verified Benchmark for Speech Emotion Detection
by: Schuhmann, Christoph, et al.
Published: (2025)
by: Schuhmann, Christoph, et al.
Published: (2025)
Stroke Disease Classification Using Machine Learning with Feature Selection Techniques
by: Hasan, Mahade, et al.
Published: (2025)
by: Hasan, Mahade, et al.
Published: (2025)
Reproducibility in Machine Learning-based Research: Overview, Barriers and Drivers
by: Semmelrock, Harald, et al.
Published: (2024)
by: Semmelrock, Harald, et al.
Published: (2024)
More Rigorous Software Engineering Would Improve Reproducibility in Machine Learning Research
by: Wolter, Moritz, et al.
Published: (2025)
by: Wolter, Moritz, et al.
Published: (2025)
What is Reproducibility in Artificial Intelligence and Machine Learning Research?
by: Desai, Abhyuday, et al.
Published: (2024)
by: Desai, Abhyuday, et al.
Published: (2024)
EmoNet-Face: An Expert-Annotated Benchmark for Synthetic Emotion Recognition
by: Schuhmann, Christoph, et al.
Published: (2025)
by: Schuhmann, Christoph, et al.
Published: (2025)
Machine Learning for Medicine Must Be Interpretable, Shareable, Reproducible and Accountable by Design
by: Bektaş, Ayyüce Begüm, et al.
Published: (2025)
by: Bektaş, Ayyüce Begüm, et al.
Published: (2025)
Automated Modernization of Machine Learning Engineering Notebooks for Reproducibility
by: Jin, Bihui, et al.
Published: (2026)
by: Jin, Bihui, et al.
Published: (2026)
Causal Representation Learning on High-Dimensional Data: Benchmarks, Reproducibility, and Evaluation Metrics
by: Sadeghi, Alireza, et al.
Published: (2026)
by: Sadeghi, Alireza, et al.
Published: (2026)
EDALearn: A Comprehensive RTL-to-Signoff EDA Benchmark for Democratized and Reproducible ML for EDA Research
by: Pan, Jingyu, et al.
Published: (2023)
by: Pan, Jingyu, et al.
Published: (2023)
Reproducibility of Machine Learning-Based Fault Detection and Diagnosis for HVAC Systems in Buildings: An Empirical Study
by: Mukhtar, Adil, et al.
Published: (2025)
by: Mukhtar, Adil, et al.
Published: (2025)
Bencher: Simple and Reproducible Benchmarking for Black-Box Optimization
by: Papenmeier, Leonard, et al.
Published: (2025)
by: Papenmeier, Leonard, et al.
Published: (2025)
Meta-Continual Mobility Forecasting for Proactive Handover Prediction
by: Mandapati, Sasi Vardhan Reddy
Published: (2025)
by: Mandapati, Sasi Vardhan Reddy
Published: (2025)
Cross-Vendor Reproducibility of Radiomics-based Machine Learning Models for Computer-aided Diagnosis
by: Chaudhary, Jatin, et al.
Published: (2024)
by: Chaudhary, Jatin, et al.
Published: (2024)
OntoAligner Meets Knowledge Graph Embedding Aligners
by: Giglou, Hamed Babaei, et al.
Published: (2025)
by: Giglou, Hamed Babaei, et al.
Published: (2025)
Machine Learning Meets Transparency in Osteoporosis Risk Assessment: A Comparative Study of ML and Explainability Analysis
by: Elias, Farhana, et al.
Published: (2025)
by: Elias, Farhana, et al.
Published: (2025)
Experimental Comparison of Light-Weight and Deep CNN Models Across Diverse Datasets
by: Papon, Md. Hefzul Hossain, et al.
Published: (2026)
by: Papon, Md. Hefzul Hossain, et al.
Published: (2026)
Transformer Based Machine Fault Detection From Audio Input
by: Holla, Kiran Voderhobli
Published: (2026)
by: Holla, Kiran Voderhobli
Published: (2026)
Autonomous Curriculum Design via Relative Entropy Based Task Modifications
by: Satici, Muhammed Yusuf, et al.
Published: (2025)
by: Satici, Muhammed Yusuf, et al.
Published: (2025)
Audio Deepfake Detection with Half-Truth Localisation Using Cross-Attentive Feature Fusion
by: Sutharya, S., et al.
Published: (2026)
by: Sutharya, S., et al.
Published: (2026)
Limiting Network Size within Finite Bounds for Optimization
by: Pinto, Linu, et al.
Published: (2019)
by: Pinto, Linu, et al.
Published: (2019)
Comparative Evaluation of Weather Forecasting using Machine Learning Models
by: Rahman, Md Saydur, et al.
Published: (2024)
by: Rahman, Md Saydur, et al.
Published: (2024)
Data Attribution in Adaptive Learning
by: Rege, Amit Kiran
Published: (2026)
by: Rege, Amit Kiran
Published: (2026)
Dynamic Configuration of On-Street Parking Spaces using Multi Agent Reinforcement Learning
by: Jayasinghe, Oshada, et al.
Published: (2025)
by: Jayasinghe, Oshada, et al.
Published: (2025)
STRABLE: Benchmarking Tabular Machine Learning with Strings
by: Blayer, Gioia, et al.
Published: (2026)
by: Blayer, Gioia, et al.
Published: (2026)
From Top-1 to Top-K: A Reproducibility Study and Benchmarking of Counterfactual Explanations for Recommender Systems
by: Nguyen, Quang-Huy, et al.
Published: (2026)
by: Nguyen, Quang-Huy, et al.
Published: (2026)
The Elusive Pursuit of Reproducing PATE-GAN: Benchmarking, Auditing, Debugging
by: Ganev, Georgi, et al.
Published: (2024)
by: Ganev, Georgi, et al.
Published: (2024)
Discussing the Spectrum of Physics-Enhanced Machine Learning; a Survey on Structural Mechanics Applications
by: Haywood-Alexander, Marcus, et al.
Published: (2023)
by: Haywood-Alexander, Marcus, et al.
Published: (2023)
Similar Items
-
AISSISTANT: Human-AI Collaborative Review and Perspective Research Workflows in Data Science
by: Gaddipati, Sasi Kiran, et al.
Published: (2025) -
MC-NEST: Enhancing Mathematical Reasoning in Large Language Models leveraging a Monte Carlo Self-Refine Tree
by: Rabby, Gollam, et al.
Published: (2024) -
SelfCheck-Eval: A Multi-Module Framework for Zero-Resource Hallucination Detection in Large Language Models
by: Muhammed, Diyana, et al.
Published: (2025) -
Iterative Hypothesis Generation for Scientific Discovery with Monte Carlo Nash Equilibrium Self-Refining Trees
by: Rabby, Gollam, et al.
Published: (2025) -
Fine-tuning and Prompt Engineering with Cognitive Knowledge Graphs for Scholarly Knowledge Organization
by: Rabby, Gollam, et al.
Published: (2024)