Towards Evaluation Engineering: An Empirical Study of ML Evaluation Harnesses in the Wild
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Zhimin, Wang, Zehao, Bangash, Abdul Ali, Adams, Bram, Hassan, Ahmed E. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
An Empirical Study of Challenges in Machine Learning Asset Management
by: Zhao, Zhimin, et al.
Published: (2024)
by: Zhao, Zhimin, et al.
Published: (2024)
On the Workflows and Smells of Leaderboard Operations (LBOps): An Exploratory Study of Foundation Model Leaderboards
by: Zhao, Zhimin, et al.
Published: (2024)
by: Zhao, Zhimin, et al.
Published: (2024)
On the Impact of Black-box Deployment Strategies for Edge AI on Latency and Model Performance
by: Singh, Jaskirat, et al.
Published: (2024)
by: Singh, Jaskirat, et al.
Published: (2024)
Understanding Prompt Management in GitHub Repositories: A Call for Best Practices
by: Li, Hao, et al.
Published: (2025)
by: Li, Hao, et al.
Published: (2025)
The State of Documentation Practices of Third-party Machine Learning Models and Datasets
by: Oreamuno, Ernesto Lang, et al.
Published: (2023)
by: Oreamuno, Ernesto Lang, et al.
Published: (2023)
Permissive-Washing in the Open AI Supply Chain: A Large-Scale Audit of License Integrity
by: Jewitt, James, et al.
Published: (2026)
by: Jewitt, James, et al.
Published: (2026)
Towards a Classification of Open-Source ML Models and Datasets for Software Engineering
by: González, Alexandra, et al.
Published: (2024)
by: González, Alexandra, et al.
Published: (2024)
Towards Semantic Versioning of Open Pre-trained Language Model Releases on Hugging Face
by: Ajibode, Adekunle, et al.
Published: (2024)
by: Ajibode, Adekunle, et al.
Published: (2024)
How Robust are LLM-Generated Library Imports? An Empirical Study using Stack Overflow
by: Latendresse, Jasmine, et al.
Published: (2025)
by: Latendresse, Jasmine, et al.
Published: (2025)
A Framework to Model ML Engineering Processes
by: Morales, Sergio, et al.
Published: (2024)
by: Morales, Sergio, et al.
Published: (2024)
HAFix: History-Augmented Large Language Models for Bug Fixing
by: Shi, Yu, et al.
Published: (2025)
by: Shi, Yu, et al.
Published: (2025)
SWE-Arena: An Interactive Platform for Evaluating Foundation Models in Software Engineering
by: Zhao, Zhimin
Published: (2025)
by: Zhao, Zhimin
Published: (2025)
A Comprehensive Framework for Evaluating API-oriented Code Generation in Large Language Models
by: Wu, Yixi, et al.
Published: (2024)
by: Wu, Yixi, et al.
Published: (2024)
Evaluating the Use of LLMs for Documentation to Code Traceability
by: Alor, Ebube, et al.
Published: (2025)
by: Alor, Ebube, et al.
Published: (2025)
An Empirical Evaluation of Locally Deployed LLMs for Bug Detection in Python Code
by: Vulićević, Jelena Ilić
Published: (2026)
by: Vulićević, Jelena Ilić
Published: (2026)
The Rise of AI Teammates in Software Engineering (SE) 3.0: How Autonomous Coding Agents Are Reshaping Software Engineering
by: Li, Hao, et al.
Published: (2025)
by: Li, Hao, et al.
Published: (2025)
Gradient-Based Model Fingerprinting for LLM Similarity Detection and Family Classification
by: Wu, Zehao, et al.
Published: (2025)
by: Wu, Zehao, et al.
Published: (2025)
HAFixAgent: History-Aware Program Repair Agent
by: Shi, Yu, et al.
Published: (2025)
by: Shi, Yu, et al.
Published: (2025)
Understanding the Helpfulness of Stale Bot for Pull-based Development: An Empirical Study of 20 Large Open-Source Projects
by: Khatoonabadi, SayedHassan, et al.
Published: (2023)
by: Khatoonabadi, SayedHassan, et al.
Published: (2023)
On the synchronization between Hugging Face pre-trained language models and their upstream GitHub repository
by: Ajibode, Adekunle, et al.
Published: (2025)
by: Ajibode, Adekunle, et al.
Published: (2025)
The State of the SBOM Tool Ecosystems: A Comparative Analysis of SPDX and CycloneDX
by: Bangash, Abdul Ali, et al.
Published: (2025)
by: Bangash, Abdul Ali, et al.
Published: (2025)
Beyond Synthetic Benchmarks: Evaluating LLM Performance on Real-World Class-Level Code Generation
by: Rahman, Musfiqur, et al.
Published: (2025)
by: Rahman, Musfiqur, et al.
Published: (2025)
OmniLLP: Enhancing LLM-based Log Level Prediction with Context-Aware Retrieval
by: Ouatiti, Youssef Esseddiq, et al.
Published: (2025)
by: Ouatiti, Youssef Esseddiq, et al.
Published: (2025)
Agentic Software Engineering: Foundational Pillars and a Research Roadmap
by: Hassan, Ahmed E., et al.
Published: (2025)
by: Hassan, Ahmed E., et al.
Published: (2025)
MermaidSeqBench: An Evaluation Benchmark for NL-to-Mermaid Sequence Diagram Generation
by: Shbita, Basel, et al.
Published: (2025)
by: Shbita, Basel, et al.
Published: (2025)
From Hazard Identification to Controller Design: Proactive and LLM-Supported Safety Engineering for ML-Powered Systems
by: Hong, Yining, et al.
Published: (2025)
by: Hong, Yining, et al.
Published: (2025)
A State-of-the-practice Release-readiness Checklist for Generative AI-based Software Products
by: Patel, Harsh, et al.
Published: (2024)
by: Patel, Harsh, et al.
Published: (2024)
A Large-Scale Study of Model Integration in ML-Enabled Software Systems
by: Sens, Yorick, et al.
Published: (2024)
by: Sens, Yorick, et al.
Published: (2024)
An Empirical Study of Fault Localisation Techniques for Deep Learning
by: Humbatova, Nargiz, et al.
Published: (2024)
by: Humbatova, Nargiz, et al.
Published: (2024)
Challenges and Paths Towards AI for Software Engineering
by: Gu, Alex, et al.
Published: (2025)
by: Gu, Alex, et al.
Published: (2025)
ML-On-Rails: Safeguarding Machine Learning Models in Software Systems A Case Study
by: Abdelkader, Hala, et al.
Published: (2024)
by: Abdelkader, Hala, et al.
Published: (2024)
On the Impact of Code Comments for Automated Bug-Fixing: An Empirical Study
by: Vitale, Antonio, et al.
Published: (2026)
by: Vitale, Antonio, et al.
Published: (2026)
Property-Driven Evaluation of GNN Expressiveness at Scale: Datasets, Framework, and Study
by: Che, Sicong, et al.
Published: (2026)
by: Che, Sicong, et al.
Published: (2026)
Beyond Output Correctness: Benchmarking and Evaluating Large Language Model Reasoning in Coding Tasks
by: Li, Yuangang, et al.
Published: (2026)
by: Li, Yuangang, et al.
Published: (2026)
From Hugging Face to GitHub: Tracing License Drift in the Open-Source AI Ecosystem
by: Jewitt, James, et al.
Published: (2025)
by: Jewitt, James, et al.
Published: (2025)
Data Quality Antipatterns for Software Analytics
by: Bhatia, Aaditya, et al.
Published: (2024)
by: Bhatia, Aaditya, et al.
Published: (2024)
Output Format Biases in the Evaluation of Large Language Models for Code Translation
by: Macedo, Marcos, et al.
Published: (2024)
by: Macedo, Marcos, et al.
Published: (2024)
Toward Explaining Large Language Models in Software Engineering Tasks
by: Vitale, Antonio, et al.
Published: (2025)
by: Vitale, Antonio, et al.
Published: (2025)
More with Less: An Empirical Study of Turn-Control Strategies for Efficient Coding Agents
by: Gao, Pengfei, et al.
Published: (2025)
by: Gao, Pengfei, et al.
Published: (2025)
Analyzing the Evolution and Maintenance of ML Models on Hugging Face
by: Castaño, Joel, et al.
Published: (2023)
by: Castaño, Joel, et al.
Published: (2023)
Similar Items
-
An Empirical Study of Challenges in Machine Learning Asset Management
by: Zhao, Zhimin, et al.
Published: (2024) -
On the Workflows and Smells of Leaderboard Operations (LBOps): An Exploratory Study of Foundation Model Leaderboards
by: Zhao, Zhimin, et al.
Published: (2024) -
On the Impact of Black-box Deployment Strategies for Edge AI on Latency and Model Performance
by: Singh, Jaskirat, et al.
Published: (2024) -
Understanding Prompt Management in GitHub Repositories: A Call for Best Practices
by: Li, Hao, et al.
Published: (2025) -
The State of Documentation Practices of Third-party Machine Learning Models and Datasets
by: Oreamuno, Ernesto Lang, et al.
Published: (2023)