Beyond Synthetic Benchmarks: Evaluating LLM Performance on Real-World Class-Level Code Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Rahman, Musfiqur, Khatoonabadi, SayedHassan, Shihab, Emad |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
OpenClassGen: A Large-Scale Corpus of Real-World Python Classes for LLM Research
by: Rahman, Musfiqur, et al.
Published: (2025)
by: Rahman, Musfiqur, et al.
Published: (2025)
Automatic Detection of LLM-Generated Code: A Comparative Case Study of Contemporary Models Across Function and Class Granularities
by: Rahman, Musfiqur, et al.
Published: (2024)
by: Rahman, Musfiqur, et al.
Published: (2024)
Evaluating the Use of LLMs for Documentation to Code Traceability
by: Alor, Ebube, et al.
Published: (2025)
by: Alor, Ebube, et al.
Published: (2025)
Evaluating the Use of LLMs for Automated DOM-Level Resolution of Web Performance Issues
by: Peters, Gideon, et al.
Published: (2026)
by: Peters, Gideon, et al.
Published: (2026)
The Impact of Environment Configurations on the Stability of AI-Enabled Systems
by: Rahman, Musfiqur, et al.
Published: (2024)
by: Rahman, Musfiqur, et al.
Published: (2024)
How Robust are LLM-Generated Library Imports? An Empirical Study using Stack Overflow
by: Latendresse, Jasmine, et al.
Published: (2025)
by: Latendresse, Jasmine, et al.
Published: (2025)
Automated File-Level Logging Generation for Machine Learning Applications using LLMs: A Case Study using GPT-4o Mini
by: Rodriguez, Mayra Sofia Ruiz, et al.
Published: (2025)
by: Rodriguez, Mayra Sofia Ruiz, et al.
Published: (2025)
Synergizing LLMs and Knowledge Graphs: A Novel Approach to Software Repository-Related Question Answering
by: Abedu, Samuel, et al.
Published: (2024)
by: Abedu, Samuel, et al.
Published: (2024)
Is ChatGPT a Good Software Librarian? An Exploratory Study on the Use of ChatGPT for Software Library Recommendations
by: Latendresse, Jasmine, et al.
Published: (2024)
by: Latendresse, Jasmine, et al.
Published: (2024)
An Approach for Auto Generation of Labeling Functions for Software Engineering Chatbots
by: Alor, Ebube, et al.
Published: (2024)
by: Alor, Ebube, et al.
Published: (2024)
On Wasted Contributions: Understanding the Dynamics of Contributor-Abandoned Pull Requests
by: Khatoonabadi, SayedHassan, et al.
Published: (2021)
by: Khatoonabadi, SayedHassan, et al.
Published: (2021)
Predicting the First Response Latency of Maintainers and Contributors in Pull Requests
by: Khatoonabadi, SayedHassan, et al.
Published: (2023)
by: Khatoonabadi, SayedHassan, et al.
Published: (2023)
Understanding the Helpfulness of Stale Bot for Pull-based Development: An Empirical Study of 20 Large Open-Source Projects
by: Khatoonabadi, SayedHassan, et al.
Published: (2023)
by: Khatoonabadi, SayedHassan, et al.
Published: (2023)
Will It Survive? Deciphering the Fate of AI-Generated Code in Open Source
by: Rahman, Musfiqur, et al.
Published: (2026)
by: Rahman, Musfiqur, et al.
Published: (2026)
Tokenomics: Quantifying Where Tokens Are Used in Agentic Software Engineering
by: Salim, Mohamad, et al.
Published: (2026)
by: Salim, Mohamad, et al.
Published: (2026)
The Impact of Large Language Models (LLMs) on Code Review Process
by: Collante, Antonio, et al.
Published: (2025)
by: Collante, Antonio, et al.
Published: (2025)
Relative Positioning Based Code Chunking Method For Rich Context Retrieval In Repository Level Code Completion Task With Code Language Model
by: Rahman, Imranur, et al.
Published: (2025)
by: Rahman, Imranur, et al.
Published: (2025)
LLM Benchmarking with LLaMA2: Evaluating Code Development Performance Across Multiple Programming Languages
by: Diehl, Patrick, et al.
Published: (2025)
by: Diehl, Patrick, et al.
Published: (2025)
From Laboratory to Real-World Applications: Benchmarking Agentic Code Reasoning at the Repository Level
by: Li, Jia, et al.
Published: (2026)
by: Li, Jia, et al.
Published: (2026)
On the Impact of Black-box Deployment Strategies for Edge AI on Latency and Model Performance
by: Singh, Jaskirat, et al.
Published: (2024)
by: Singh, Jaskirat, et al.
Published: (2024)
Beyond Output Correctness: Benchmarking and Evaluating Large Language Model Reasoning in Coding Tasks
by: Li, Yuangang, et al.
Published: (2026)
by: Li, Yuangang, et al.
Published: (2026)
Unseen Horizons: Unveiling the Real Capability of LLM Code Generation Beyond the Familiar
by: Zhang, Yuanliang, et al.
Published: (2024)
by: Zhang, Yuanliang, et al.
Published: (2024)
How Efficient is LLM-Generated Code? A Rigorous & High-Standard Benchmark
by: Qiu, Ruizhong, et al.
Published: (2024)
by: Qiu, Ruizhong, et al.
Published: (2024)
CodeSense: a Real-World Benchmark and Dataset for Code Semantic Reasoning
by: Roy, Monoshi Kumar, et al.
Published: (2025)
by: Roy, Monoshi Kumar, et al.
Published: (2025)
Beyond Code Snippets: Benchmarking LLMs on Repository-Level Question Answering
by: Alebachew, Yoseph Berhanu, et al.
Published: (2026)
by: Alebachew, Yoseph Berhanu, et al.
Published: (2026)
A.S.E: A Repository-Level Benchmark for Evaluating Security in AI-Generated Code
by: Lian, Keke, et al.
Published: (2025)
by: Lian, Keke, et al.
Published: (2025)
Automated Benchmark Generation for Repository-Level Coding Tasks
by: Vergopoulos, Konstantinos, et al.
Published: (2025)
by: Vergopoulos, Konstantinos, et al.
Published: (2025)
CodeTaste: Can LLMs Generate Human-Level Code Refactorings?
by: Thillen, Alex, et al.
Published: (2026)
by: Thillen, Alex, et al.
Published: (2026)
1D-Bench: A Benchmark for Iterative UI Code Generation with Visual Feedback in Real-World
by: Xu, Qiao, et al.
Published: (2026)
by: Xu, Qiao, et al.
Published: (2026)
Uncovering LLM-Generated Code: A Zero-Shot Synthetic Code Detector via Code Rewriting
by: Ye, Tong, et al.
Published: (2024)
by: Ye, Tong, et al.
Published: (2024)
EvoCodeBench: An Evolving Code Generation Benchmark Aligned with Real-World Code Repositories
by: Li, Jia, et al.
Published: (2024)
by: Li, Jia, et al.
Published: (2024)
SWT-Bench: Testing and Validating Real-World Bug-Fixes with Code Agents
by: Mündler, Niels, et al.
Published: (2024)
by: Mündler, Niels, et al.
Published: (2024)
MobiFlow: Real-World Mobile Agent Benchmarking through Trajectory Fusion
by: Feng, Yunfei, et al.
Published: (2026)
by: Feng, Yunfei, et al.
Published: (2026)
Re-Evaluating Code LLM Benchmarks Under Semantic Mutation
by: Pan, Zhiyuan, et al.
Published: (2025)
by: Pan, Zhiyuan, et al.
Published: (2025)
Mage: Multi-Axis Evaluation of LLM-Generated Executable Game Scenes Beyond Compile-Pass Rate
by: Liu, Hugh Xuechen, et al.
Published: (2026)
by: Liu, Hugh Xuechen, et al.
Published: (2026)
RepoTransBench: A Real-World Multilingual Benchmark for Repository-Level Code Translation
by: Wang, Yanli, et al.
Published: (2024)
by: Wang, Yanli, et al.
Published: (2024)
Beyond Functional Correctness: Exploring Hallucinations in LLM-Generated Code
by: Liu, Fang, et al.
Published: (2024)
by: Liu, Fang, et al.
Published: (2024)
Deep-Bench: Deep Learning Benchmark Dataset for Code Generation
by: Daghighfarsoodeh, Alireza, et al.
Published: (2025)
by: Daghighfarsoodeh, Alireza, et al.
Published: (2025)
Drawing Pandas: A Benchmark for LLMs in Generating Plotting Code
by: Galimzyanov, Timur, et al.
Published: (2024)
by: Galimzyanov, Timur, et al.
Published: (2024)
GitTaskBench: A Benchmark for Code Agents Solving Real-World Tasks Through Code Repository Leveraging
by: Ni, Ziyi, et al.
Published: (2025)
by: Ni, Ziyi, et al.
Published: (2025)
Similar Items
-
OpenClassGen: A Large-Scale Corpus of Real-World Python Classes for LLM Research
by: Rahman, Musfiqur, et al.
Published: (2025) -
Automatic Detection of LLM-Generated Code: A Comparative Case Study of Contemporary Models Across Function and Class Granularities
by: Rahman, Musfiqur, et al.
Published: (2024) -
Evaluating the Use of LLMs for Documentation to Code Traceability
by: Alor, Ebube, et al.
Published: (2025) -
Evaluating the Use of LLMs for Automated DOM-Level Resolution of Web Performance Issues
by: Peters, Gideon, et al.
Published: (2026) -
The Impact of Environment Configurations on the Stability of AI-Enabled Systems
by: Rahman, Musfiqur, et al.
Published: (2024)