On the Workflows and Smells of Leaderboard Operations (LBOps): An Exploratory Study of Foundation Model Leaderboards
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Zhimin, Bangash, Abdul Ali, Côgo, Filipe Roseiro, Adams, Bram, Hassan, Ahmed E. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Semantic Versioning of Open Pre-trained Language Model Releases on Hugging Face
by: Ajibode, Adekunle, et al.
Published: (2024)
by: Ajibode, Adekunle, et al.
Published: (2024)
Towards Evaluation Engineering: An Empirical Study of ML Evaluation Harnesses in the Wild
by: Zhao, Zhimin, et al.
Published: (2026)
by: Zhao, Zhimin, et al.
Published: (2026)
An Empirical Study of Challenges in Machine Learning Asset Management
by: Zhao, Zhimin, et al.
Published: (2024)
by: Zhao, Zhimin, et al.
Published: (2024)
Understanding Prompt Management in GitHub Repositories: A Call for Best Practices
by: Li, Hao, et al.
Published: (2025)
by: Li, Hao, et al.
Published: (2025)
From Leaderboard to Deployment: Code Quality Challenges in AV Perception Repositories
by: Karvat, Mateus, et al.
Published: (2026)
by: Karvat, Mateus, et al.
Published: (2026)
HAFix: History-Augmented Large Language Models for Bug Fixing
by: Shi, Yu, et al.
Published: (2025)
by: Shi, Yu, et al.
Published: (2025)
On the synchronization between Hugging Face pre-trained language models and their upstream GitHub repository
by: Ajibode, Adekunle, et al.
Published: (2025)
by: Ajibode, Adekunle, et al.
Published: (2025)
RepairBench: Leaderboard of Frontier Models for Program Repair
by: Silva, André, et al.
Published: (2024)
by: Silva, André, et al.
Published: (2024)
The State of the SBOM Tool Ecosystems: A Comparative Analysis of SPDX and CycloneDX
by: Bangash, Abdul Ali, et al.
Published: (2025)
by: Bangash, Abdul Ali, et al.
Published: (2025)
The State of Documentation Practices of Third-party Machine Learning Models and Datasets
by: Oreamuno, Ernesto Lang, et al.
Published: (2023)
by: Oreamuno, Ernesto Lang, et al.
Published: (2023)
Leveraging the Crowd for Dependency Management: An Empirical Study on the Dependabot Compatibility Score
by: Rombaut, Benjamin, et al.
Published: (2024)
by: Rombaut, Benjamin, et al.
Published: (2024)
Output Format Biases in the Evaluation of Large Language Models for Code Translation
by: Macedo, Marcos, et al.
Published: (2024)
by: Macedo, Marcos, et al.
Published: (2024)
Continuous Integration Practices in Machine Learning Projects: The Practitioners` Perspective
by: Bernardo, João Helis, et al.
Published: (2025)
by: Bernardo, João Helis, et al.
Published: (2025)
On the Impact of Black-box Deployment Strategies for Edge AI on Latency and Model Performance
by: Singh, Jaskirat, et al.
Published: (2024)
by: Singh, Jaskirat, et al.
Published: (2024)
Compiler.next: A Search-Based Compiler to Power the AI-Native Future of Software Engineering
by: Cogo, Filipe R., et al.
Published: (2025)
by: Cogo, Filipe R., et al.
Published: (2025)
A Large-Scale Exploratory Study on the Proxy Pattern in Ethereum
by: Ebrahimi, Amir M., et al.
Published: (2025)
by: Ebrahimi, Amir M., et al.
Published: (2025)
SWE-Arena: An Interactive Platform for Evaluating Foundation Models in Software Engineering
by: Zhao, Zhimin
Published: (2025)
by: Zhao, Zhimin
Published: (2025)
Does Using Bazel Help Speed Up Continuous Integration Builds?
by: Zheng, Shenyu, et al.
Published: (2024)
by: Zheng, Shenyu, et al.
Published: (2024)
An Empirical Study of Self-Admitted Technical Debt in Machine Learning Software
by: Bhatia, Aaditya, et al.
Published: (2023)
by: Bhatia, Aaditya, et al.
Published: (2023)
I depended on you and you broke me: An empirical study of manifesting breaking changes in client packages
by: Venturini, Daniel, et al.
Published: (2023)
by: Venturini, Daniel, et al.
Published: (2023)
InterTrans: Leveraging Transitive Intermediate Translations to Enhance LLM-based Code Translation
by: Macedo, Marcos, et al.
Published: (2024)
by: Macedo, Marcos, et al.
Published: (2024)
An Empirical Study of Developers' Challenges in Implementing Workflows as Code: A Case Study on Apache Airflow
by: Yasmin, Jerin, et al.
Published: (2024)
by: Yasmin, Jerin, et al.
Published: (2024)
A Comprehensive Evaluation of Four End-to-End AI Autopilots Using CCTest and the Carla Leaderboard
by: Li, Changwen, et al.
Published: (2025)
by: Li, Changwen, et al.
Published: (2025)
AgenticSZZ: Temporal Knowledge Graph-Guided Agentic Bug-Inducing Commit Identification
by: Shi, Yu, et al.
Published: (2026)
by: Shi, Yu, et al.
Published: (2026)
Top Leaderboard Ranking = Top Coding Proficiency, Always? EvoEval: Evolving Coding Benchmarks via LLM
by: Xia, Chunqiu Steven, et al.
Published: (2024)
by: Xia, Chunqiu Steven, et al.
Published: (2024)
Assessing and Improving the Representativeness of Code Generation Benchmarks Using Knowledge Units (KUs) of Programming Languages -- An Empirical Study
by: Ahasanuzzaman, Md, et al.
Published: (2026)
by: Ahasanuzzaman, Md, et al.
Published: (2026)
The Hitchhikers Guide to Production-ready Trustworthy Foundation Model powered Software (FMware)
by: Vasilevski, Kirill, et al.
Published: (2025)
by: Vasilevski, Kirill, et al.
Published: (2025)
IRJIT: A Simple, Online, Information Retrieval Approach for Just-In-Time Software Defect Prediction
by: Sahar, Hareem, et al.
Published: (2022)
by: Sahar, Hareem, et al.
Published: (2022)
AgentAtlas: Beyond Outcome Leaderboards for LLM Agents
by: Mazaheri, Parsa, et al.
Published: (2026)
by: Mazaheri, Parsa, et al.
Published: (2026)
Towards Reliable Generation of Executable Workflows by Foundation Models
by: Masoumzadeh, Sogol, et al.
Published: (2025)
by: Masoumzadeh, Sogol, et al.
Published: (2025)
HAFixAgent: History-Aware Program Repair Agent
by: Shi, Yu, et al.
Published: (2025)
by: Shi, Yu, et al.
Published: (2025)
Permissive-Washing in the Open AI Supply Chain: A Large-Scale Audit of License Integrity
by: Jewitt, James, et al.
Published: (2026)
by: Jewitt, James, et al.
Published: (2026)
Agentic Software Engineering: Foundational Pillars and a Research Roadmap
by: Hassan, Ahmed E., et al.
Published: (2025)
by: Hassan, Ahmed E., et al.
Published: (2025)
Reliability of AI Bots Footprints in GitHub Actions CI/CD Workflows
by: Shah, Syed Muhammad Ashhar, et al.
Published: (2026)
by: Shah, Syed Muhammad Ashhar, et al.
Published: (2026)
OmniLLP: Enhancing LLM-based Log Level Prediction with Context-Aware Retrieval
by: Ouatiti, Youssef Esseddiq, et al.
Published: (2025)
by: Ouatiti, Youssef Esseddiq, et al.
Published: (2025)
Dissecting the SWE-Bench Leaderboards: Profiling Submitters and Architectures of LLM- and Agent-Based Repair Systems
by: Martinez, Matias, et al.
Published: (2025)
by: Martinez, Matias, et al.
Published: (2025)
Model Context Protocol (MCP) at First Glance: Studying the Security and Maintainability of MCP Servers
by: Hasan, Mohammed Mehedi, et al.
Published: (2025)
by: Hasan, Mohammed Mehedi, et al.
Published: (2025)
Novice Developers Produce Larger Review Overhead for Project Maintainers while Vibe Coding
by: Asdaque, Syed Ammar, et al.
Published: (2026)
by: Asdaque, Syed Ammar, et al.
Published: (2026)
An Empirical Study on Code Review Activity Prediction and Its Impact in Practice
by: Olewicki, Doriane, et al.
Published: (2024)
by: Olewicki, Doriane, et al.
Published: (2024)
Comparative Analysis of Quantum and Classical Support Vector Classifiers for Software Bug Prediction: An Exploratory Study
by: Nadim, Md, et al.
Published: (2025)
by: Nadim, Md, et al.
Published: (2025)
Similar Items
-
Towards Semantic Versioning of Open Pre-trained Language Model Releases on Hugging Face
by: Ajibode, Adekunle, et al.
Published: (2024) -
Towards Evaluation Engineering: An Empirical Study of ML Evaluation Harnesses in the Wild
by: Zhao, Zhimin, et al.
Published: (2026) -
An Empirical Study of Challenges in Machine Learning Asset Management
by: Zhao, Zhimin, et al.
Published: (2024) -
Understanding Prompt Management in GitHub Repositories: A Call for Best Practices
by: Li, Hao, et al.
Published: (2025) -
From Leaderboard to Deployment: Code Quality Challenges in AV Perception Repositories
by: Karvat, Mateus, et al.
Published: (2026)