Design Principles for Falsifiable, Replicable and Reproducible Empirical ML Research
Fuente:
arXiv
Saved in:
| Main Authors: | Vranješ, Daniel, Niggemann, Oliver |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Data Driven Diagnosis for Large Cyber-Physical-Systems with Minimal Prior Information
by: Steude, Henrik Sebastian, et al.
Published: (2025)
by: Steude, Henrik Sebastian, et al.
Published: (2025)
Quantifying Robustness: A Benchmarking Framework for Deep Learning Forecasting in Cyber-Physical Systems
by: Windmann, Alexander, et al.
Published: (2025)
by: Windmann, Alexander, et al.
Published: (2025)
MAWIFlow Benchmark: Realistic Flow-Based Evaluation for Network Intrusion Detection
by: Schraven, Joshua, et al.
Published: (2025)
by: Schraven, Joshua, et al.
Published: (2025)
On the Replicability and Reproducibility of Deep Learning in Software Engineering
by: Liu, Chao, et al.
Published: (2020)
by: Liu, Chao, et al.
Published: (2020)
Unlearning as Ablation: Toward a Falsifiable Benchmark for Generative Scientific Discovery
by: Yang, Robert
Published: (2025)
by: Yang, Robert
Published: (2025)
Grokking as a Falsifiable Finite-Size Transition
by: Bi, Yuda, et al.
Published: (2026)
by: Bi, Yuda, et al.
Published: (2026)
A Reproducible Log-Driven AutoML Framework for Interpretable Pipeline Optimization in Healthcare Risk Prediction
by: Huang, Rui, et al.
Published: (2026)
by: Huang, Rui, et al.
Published: (2026)
Avionic Main Fuel Pump Simulation and Fault-Diagnosis Benchmark
by: Janzen, Felix Leonhard, et al.
Published: (2026)
by: Janzen, Felix Leonhard, et al.
Published: (2026)
Artificial Intelligence in Industry 4.0: A Review of Integration Challenges for Industrial Systems
by: Windmann, Alexander, et al.
Published: (2024)
by: Windmann, Alexander, et al.
Published: (2024)
Reproducibility of Machine Learning-Based Fault Detection and Diagnosis for HVAC Systems in Buildings: An Empirical Study
by: Mukhtar, Adil, et al.
Published: (2025)
by: Mukhtar, Adil, et al.
Published: (2025)
Kolmogorov-Arnold Convolutions: Design Principles and Empirical Studies
by: Drokin, Ivan
Published: (2024)
by: Drokin, Ivan
Published: (2024)
Can LLMs Revolutionize the Design of Explainable and Efficient TinyML Models?
by: Zeinaty, Christophe El, et al.
Published: (2025)
by: Zeinaty, Christophe El, et al.
Published: (2025)
Forest vs Tree: The $(N, K)$ Trade-off in Reproducible ML Evaluation
by: Pandita, Deepak, et al.
Published: (2025)
by: Pandita, Deepak, et al.
Published: (2025)
What Do Machine Learning Researchers Mean by "Reproducible"?
by: Raff, Edward, et al.
Published: (2024)
by: Raff, Edward, et al.
Published: (2024)
Learning to be Reproducible: Custom Loss Design for Robust Neural Networks
by: Ahmed, Waqas, et al.
Published: (2026)
by: Ahmed, Waqas, et al.
Published: (2026)
Empirical Design in Reinforcement Learning
by: Patterson, Andrew, et al.
Published: (2023)
by: Patterson, Andrew, et al.
Published: (2023)
Empirical Evaluation of Concept Drift in ML-Based Android Malware Detection
by: Sabbah, Ahmed, et al.
Published: (2025)
by: Sabbah, Ahmed, et al.
Published: (2025)
Towards Evaluation Engineering: An Empirical Study of ML Evaluation Harnesses in the Wild
by: Zhao, Zhimin, et al.
Published: (2026)
by: Zhao, Zhimin, et al.
Published: (2026)
Towards Reproducibility in Predictive Process Mining: SPICE -- A Deep Learning Library
by: Stritzel, Oliver, et al.
Published: (2025)
by: Stritzel, Oliver, et al.
Published: (2025)
List Replicable Reinforcement Learning
by: Zhang, Bohan, et al.
Published: (2025)
by: Zhang, Bohan, et al.
Published: (2025)
Beyond GRPO and On-Policy Distillation: An Empirical Sparse-to-Dense Reward Principle for Language-Model Post-Training
by: Xu, Yuanda, et al.
Published: (2026)
by: Xu, Yuanda, et al.
Published: (2026)
Design Principles for Sequence Models via Coefficient Dynamics
by: Sieber, Jerome, et al.
Published: (2025)
by: Sieber, Jerome, et al.
Published: (2025)
Reproducibility study of FairAC
by: de Jong, Gijs, et al.
Published: (2024)
by: de Jong, Gijs, et al.
Published: (2024)
Design Once, Deploy at Scale: Template-Driven ML Development for Large Model Ecosystems
by: Liu, Jiang, et al.
Published: (2026)
by: Liu, Jiang, et al.
Published: (2026)
Towards Principled Graph Transformers
by: Müller, Luis, et al.
Published: (2024)
by: Müller, Luis, et al.
Published: (2024)
ML-Tool-Bench: Tool-Augmented Planning for ML Tasks
by: Chittepu, Yaswanth, et al.
Published: (2025)
by: Chittepu, Yaswanth, et al.
Published: (2025)
ML For Hardware Design Interpretability: Challenges and Opportunities
by: Baartmans, Raymond, et al.
Published: (2025)
by: Baartmans, Raymond, et al.
Published: (2025)
Initializing Services in Interactive ML Systems for Diverse Users
by: Bose, Avinandan, et al.
Published: (2023)
by: Bose, Avinandan, et al.
Published: (2023)
Policy Newton Algorithm in Reproducing Kernel Hilbert Space
by: Zhang, Yixian, et al.
Published: (2025)
by: Zhang, Yixian, et al.
Published: (2025)
Reward Learning from Best-of-$N$ Preference Data: Targets, Tradeoffs, and Design Principles
by: Pukdee, Rattana, et al.
Published: (2026)
by: Pukdee, Rattana, et al.
Published: (2026)
Deriva-ML: A Continuous FAIRness Approach to Reproducible Machine Learning Models
by: Li, Zhiwei, et al.
Published: (2024)
by: Li, Zhiwei, et al.
Published: (2024)
Position: An Empirically Grounded Identifiability Theory Will Accelerate Self-Supervised Learning Research
by: Reizinger, Patrik, et al.
Published: (2025)
by: Reizinger, Patrik, et al.
Published: (2025)
A Falsifiable Prediction of Gravitational Wave Echoes from a Regularized Kerr-Newman Spacetime Derived via Architectural Analysis
by: Billions, Ava, et al.
Published: (2025)
by: Billions, Ava, et al.
Published: (2025)
Improving Reproducibility in Evaluation through Multi-Level Annotator Modeling
by: Pandita, Deepak, et al.
Published: (2026)
by: Pandita, Deepak, et al.
Published: (2026)
Completeness of Datasets Documentation on ML/AI repositories: an Empirical Investigation
by: Rondina, Marco, et al.
Published: (2025)
by: Rondina, Marco, et al.
Published: (2025)
Problem-oriented AutoML in Clustering
by: da Silva, Matheus Camilo, et al.
Published: (2024)
by: da Silva, Matheus Camilo, et al.
Published: (2024)
nanoML for Human Activity Recognition
by: Bacellar, Alan T. L., et al.
Published: (2025)
by: Bacellar, Alan T. L., et al.
Published: (2025)
Evaluating the printability of stl files with ML
by: Henn, Janik, et al.
Published: (2025)
by: Henn, Janik, et al.
Published: (2025)
Optimizing ML Training with Metagradient Descent
by: Engstrom, Logan, et al.
Published: (2025)
by: Engstrom, Logan, et al.
Published: (2025)
pAI/MSc: ML Theory Research with Humans on the Loop
by: Abdelmoneum, Mahmoud, et al.
Published: (2026)
by: Abdelmoneum, Mahmoud, et al.
Published: (2026)
Similar Items
-
Data Driven Diagnosis for Large Cyber-Physical-Systems with Minimal Prior Information
by: Steude, Henrik Sebastian, et al.
Published: (2025) -
Quantifying Robustness: A Benchmarking Framework for Deep Learning Forecasting in Cyber-Physical Systems
by: Windmann, Alexander, et al.
Published: (2025) -
MAWIFlow Benchmark: Realistic Flow-Based Evaluation for Network Intrusion Detection
by: Schraven, Joshua, et al.
Published: (2025) -
On the Replicability and Reproducibility of Deep Learning in Software Engineering
by: Liu, Chao, et al.
Published: (2020) -
Unlearning as Ablation: Toward a Falsifiable Benchmark for Generative Scientific Discovery
by: Yang, Robert
Published: (2025)