Assessing and Improving the Representativeness of Code Generation Benchmarks Using Knowledge Units (KUs) of Programming Languages -- An Empirical Study
Fuente:
arXiv
Saved in:
| Main Authors: | Ahasanuzzaman, Md, Adams, Bram, Fallahzadeh, Emad, Oliva, Gustavo A., Hassan, Ahmed E. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Predicting post-release defects with knowledge units (KUs) of programming languages: an empirical study
by: Ahasanuzzaman, Md, et al.
Published: (2024)
by: Ahasanuzzaman, Md, et al.
Published: (2024)
Predicting long time contributors with knowledge units of programming languages: an empirical study
by: Ahasanuzzaman, Md, et al.
Published: (2024)
by: Ahasanuzzaman, Md, et al.
Published: (2024)
HAFix: History-Augmented Large Language Models for Bug Fixing
by: Shi, Yu, et al.
Published: (2025)
by: Shi, Yu, et al.
Published: (2025)
On the Impact of Black-box Deployment Strategies for Edge AI on Latency and Model Performance
by: Singh, Jaskirat, et al.
Published: (2024)
by: Singh, Jaskirat, et al.
Published: (2024)
An Empirical Study of Testing Practices in Open Source AI Agent Frameworks and Agentic Applications
by: Hasan, Mohammed Mehedi, et al.
Published: (2025)
by: Hasan, Mohammed Mehedi, et al.
Published: (2025)
A State-of-the-practice Release-readiness Checklist for Generative AI-based Software Products
by: Patel, Harsh, et al.
Published: (2024)
by: Patel, Harsh, et al.
Published: (2024)
Model Context Protocol (MCP) at First Glance: Studying the Security and Maintainability of MCP Servers
by: Hasan, Mohammed Mehedi, et al.
Published: (2025)
by: Hasan, Mohammed Mehedi, et al.
Published: (2025)
An Empirical Study of Self-Admitted Technical Debt in Machine Learning Software
by: Bhatia, Aaditya, et al.
Published: (2023)
by: Bhatia, Aaditya, et al.
Published: (2023)
An Empirical Study on Code Review Activity Prediction and Its Impact in Practice
by: Olewicki, Doriane, et al.
Published: (2024)
by: Olewicki, Doriane, et al.
Published: (2024)
Assessing Small Language Models for Code Generation: An Empirical Study with Benchmarks
by: Hasan, Md Mahade, et al.
Published: (2025)
by: Hasan, Md Mahade, et al.
Published: (2025)
Agentic Refactoring: An Empirical Study of AI Coding Agents
by: Horikawa, Kosei, et al.
Published: (2025)
by: Horikawa, Kosei, et al.
Published: (2025)
Does Using Bazel Help Speed Up Continuous Integration Builds?
by: Zheng, Shenyu, et al.
Published: (2024)
by: Zheng, Shenyu, et al.
Published: (2024)
An Empirical Study of Developers' Challenges in Implementing Workflows as Code: A Case Study on Apache Airflow
by: Yasmin, Jerin, et al.
Published: (2024)
by: Yasmin, Jerin, et al.
Published: (2024)
Agent READMEs: An Empirical Study of Context Files for Agentic Coding
by: Chatlatanagulchai, Worawalan, et al.
Published: (2025)
by: Chatlatanagulchai, Worawalan, et al.
Published: (2025)
A Large-Scale Exploratory Study on the Proxy Pattern in Ethereum
by: Ebrahimi, Amir M., et al.
Published: (2025)
by: Ebrahimi, Amir M., et al.
Published: (2025)
AgenticSZZ: Temporal Knowledge Graph-Guided Agentic Bug-Inducing Commit Identification
by: Shi, Yu, et al.
Published: (2026)
by: Shi, Yu, et al.
Published: (2026)
An Empirical Study of Challenges in Machine Learning Asset Management
by: Zhao, Zhimin, et al.
Published: (2024)
by: Zhao, Zhimin, et al.
Published: (2024)
HAFixAgent: History-Aware Program Repair Agent
by: Shi, Yu, et al.
Published: (2025)
by: Shi, Yu, et al.
Published: (2025)
UPC Sentinel: An Accurate Approach for Detecting Upgradeability Proxy Contracts in Ethereum
by: Ebrahimi, Amir M., et al.
Published: (2024)
by: Ebrahimi, Amir M., et al.
Published: (2024)
Towards Evaluation Engineering: An Empirical Study of ML Evaluation Harnesses in the Wild
by: Zhao, Zhimin, et al.
Published: (2026)
by: Zhao, Zhimin, et al.
Published: (2026)
Developer-LLM Conversations: An Empirical Study of Interactions and Generated Code Quality
by: Zhong, Suzhen, et al.
Published: (2025)
by: Zhong, Suzhen, et al.
Published: (2025)
Do AI Coding Agents Log Like Humans? An Empirical Study
by: Ouatiti, Youssef Esseddiq, et al.
Published: (2026)
by: Ouatiti, Youssef Esseddiq, et al.
Published: (2026)
Towards Semantic Versioning of Open Pre-trained Language Model Releases on Hugging Face
by: Ajibode, Adekunle, et al.
Published: (2024)
by: Ajibode, Adekunle, et al.
Published: (2024)
On the Use of Agentic Coding: An Empirical Study of Pull Requests on GitHub
by: Watanabe, Miku, et al.
Published: (2025)
by: Watanabe, Miku, et al.
Published: (2025)
Model Compression vs. Adversarial Robustness: An Empirical Study on Language Models for Code
by: Awal, Md. Abdul, et al.
Published: (2025)
by: Awal, Md. Abdul, et al.
Published: (2025)
Human-AI Synergy in Agentic Code Review
by: Zhong, Suzhen, et al.
Published: (2026)
by: Zhong, Suzhen, et al.
Published: (2026)
On the Costs and Benefits of Adopting Lifelong Learning for Software Analytics -- Empirical Study on Brown Build and Risk Prediction
by: Olewicki, Doriane, et al.
Published: (2023)
by: Olewicki, Doriane, et al.
Published: (2023)
Model Context Protocol (MCP) Tool Descriptions Are Smelly! Towards Improving AI Agent Efficiency with Augmented MCP Tool Descriptions
by: Hasan, Mohammed Mehedi, et al.
Published: (2026)
by: Hasan, Mohammed Mehedi, et al.
Published: (2026)
Output Format Biases in the Evaluation of Large Language Models for Code Translation
by: Macedo, Marcos, et al.
Published: (2024)
by: Macedo, Marcos, et al.
Published: (2024)
On the synchronization between Hugging Face pre-trained language models and their upstream GitHub repository
by: Ajibode, Adekunle, et al.
Published: (2025)
by: Ajibode, Adekunle, et al.
Published: (2025)
OmniLLP: Enhancing LLM-based Log Level Prediction with Context-Aware Retrieval
by: Ouatiti, Youssef Esseddiq, et al.
Published: (2025)
by: Ouatiti, Youssef Esseddiq, et al.
Published: (2025)
An Empirical Study on Low-Code Programming using Traditional vs Large Language Model Support
by: Liu, Yongkun, et al.
Published: (2024)
by: Liu, Yongkun, et al.
Published: (2024)
The Impact of Large Language Models (LLMs) on Code Review Process
by: Collante, Antonio, et al.
Published: (2025)
by: Collante, Antonio, et al.
Published: (2025)
On the Workflows and Smells of Leaderboard Operations (LBOps): An Exploratory Study of Foundation Model Leaderboards
by: Zhao, Zhimin, et al.
Published: (2024)
by: Zhao, Zhimin, et al.
Published: (2024)
An Empirical Study of Python Library Migration Using Large Language Models
by: Islam, Md Mohayeminul, et al.
Published: (2025)
by: Islam, Md Mohayeminul, et al.
Published: (2025)
Compiler.next: A Search-Based Compiler to Power the AI-Native Future of Software Engineering
by: Cogo, Filipe R., et al.
Published: (2025)
by: Cogo, Filipe R., et al.
Published: (2025)
Individual Differences Limit Predicting Well-being and Productivity Using Software Repositories: A Longitudinal Industrial Study
by: Kuutila, Miikka, et al.
Published: (2021)
by: Kuutila, Miikka, et al.
Published: (2021)
Leveraging the Crowd for Dependency Management: An Empirical Study on the Dependabot Compatibility Score
by: Rombaut, Benjamin, et al.
Published: (2024)
by: Rombaut, Benjamin, et al.
Published: (2024)
Beyond Synthetic Benchmarks: Evaluating LLM Performance on Real-World Class-Level Code Generation
by: Rahman, Musfiqur, et al.
Published: (2025)
by: Rahman, Musfiqur, et al.
Published: (2025)
How Robust are LLM-Generated Library Imports? An Empirical Study using Stack Overflow
by: Latendresse, Jasmine, et al.
Published: (2025)
by: Latendresse, Jasmine, et al.
Published: (2025)
Similar Items
-
Predicting post-release defects with knowledge units (KUs) of programming languages: an empirical study
by: Ahasanuzzaman, Md, et al.
Published: (2024) -
Predicting long time contributors with knowledge units of programming languages: an empirical study
by: Ahasanuzzaman, Md, et al.
Published: (2024) -
HAFix: History-Augmented Large Language Models for Bug Fixing
by: Shi, Yu, et al.
Published: (2025) -
On the Impact of Black-box Deployment Strategies for Edge AI on Latency and Model Performance
by: Singh, Jaskirat, et al.
Published: (2024) -
An Empirical Study of Testing Practices in Open Source AI Agent Frameworks and Agentic Applications
by: Hasan, Mohammed Mehedi, et al.
Published: (2025)