OSS-Bench: Benchmark Generator for Coding LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jiang, Yuancheng, Yap, Roland, Liang, Zhenkai |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Enhanced Differential Testing in Emerging Database Systems
von: Jiang, Yuancheng, et al.
Veröffentlicht: (2025)
von: Jiang, Yuancheng, et al.
Veröffentlicht: (2025)
AICD Bench: A Challenging Benchmark for AI-Generated Code Detection
von: Orel, Daniil, et al.
Veröffentlicht: (2026)
von: Orel, Daniil, et al.
Veröffentlicht: (2026)
ThrowBench: Benchmarking LLMs by Predicting Runtime Exceptions
von: Prenner, Julian Aron, et al.
Veröffentlicht: (2025)
von: Prenner, Julian Aron, et al.
Veröffentlicht: (2025)
Deep-Bench: Deep Learning Benchmark Dataset for Code Generation
von: Daghighfarsoodeh, Alireza, et al.
Veröffentlicht: (2025)
von: Daghighfarsoodeh, Alireza, et al.
Veröffentlicht: (2025)
Where Do LLMs Still Struggle? An In-Depth Analysis of Code Generation Benchmarks
von: Sharifloo, Amir Molzam, et al.
Veröffentlicht: (2025)
von: Sharifloo, Amir Molzam, et al.
Veröffentlicht: (2025)
CONCUR: Benchmarking LLMs for Concurrent Code Generation
von: Huang, Jue, et al.
Veröffentlicht: (2026)
von: Huang, Jue, et al.
Veröffentlicht: (2026)
DevBench: A Realistic, Developer-Informed Benchmark for Code Generation Models
von: Kumarappan, Adarsh, et al.
Veröffentlicht: (2026)
von: Kumarappan, Adarsh, et al.
Veröffentlicht: (2026)
Can I Solve It? Identifying APIs Required to Complete OSS Task
von: Santos, Fabio, et al.
Veröffentlicht: (2021)
von: Santos, Fabio, et al.
Veröffentlicht: (2021)
Drawing Pandas: A Benchmark for LLMs in Generating Plotting Code
von: Galimzyanov, Timur, et al.
Veröffentlicht: (2024)
von: Galimzyanov, Timur, et al.
Veröffentlicht: (2024)
AssertionBench: A Benchmark to Evaluate Large-Language Models for Assertion Generation
von: Pulavarthi, Vaishnavi, et al.
Veröffentlicht: (2024)
von: Pulavarthi, Vaishnavi, et al.
Veröffentlicht: (2024)
CodeIF: Benchmarking the Instruction-Following Capabilities of Large Language Models for Code Generation
von: Yan, Kaiwen, et al.
Veröffentlicht: (2025)
von: Yan, Kaiwen, et al.
Veröffentlicht: (2025)
The Fault in our Stars: Quality Assessment of Code Generation Benchmarks
von: Siddiq, Mohammed Latif, et al.
Veröffentlicht: (2024)
von: Siddiq, Mohammed Latif, et al.
Veröffentlicht: (2024)
EnvBench: A Benchmark for Automated Environment Setup
von: Eliseeva, Aleksandra, et al.
Veröffentlicht: (2025)
von: Eliseeva, Aleksandra, et al.
Veröffentlicht: (2025)
StackEval: Benchmarking LLMs in Coding Assistance
von: Shah, Nidhish, et al.
Veröffentlicht: (2024)
von: Shah, Nidhish, et al.
Veröffentlicht: (2024)
Prism: Dynamic and Flexible Benchmarking of LLMs Code Generation with Monte Carlo Tree Search
von: Majdinasab, Vahid, et al.
Veröffentlicht: (2025)
von: Majdinasab, Vahid, et al.
Veröffentlicht: (2025)
Leveraging LLMs for Legacy Code Modernization: Challenges and Opportunities for LLM-Generated Documentation
von: Diggs, Colin, et al.
Veröffentlicht: (2024)
von: Diggs, Colin, et al.
Veröffentlicht: (2024)
CoDocBench: A Dataset for Code-Documentation Alignment in Software Maintenance
von: Pai, Kunal, et al.
Veröffentlicht: (2025)
von: Pai, Kunal, et al.
Veröffentlicht: (2025)
Teaching Code Refactoring Using LLMs
von: Khairnar, Anshul, et al.
Veröffentlicht: (2025)
von: Khairnar, Anshul, et al.
Veröffentlicht: (2025)
Towards Verified Code Reasoning by LLMs
von: Sistla, Meghana, et al.
Veröffentlicht: (2025)
von: Sistla, Meghana, et al.
Veröffentlicht: (2025)
InfiBench: Evaluating the Question-Answering Capabilities of Code Large Language Models
von: Li, Linyi, et al.
Veröffentlicht: (2024)
von: Li, Linyi, et al.
Veröffentlicht: (2024)
CPP-UT-Bench: Can LLMs Write Complex Unit Tests in C++?
von: Bhargava, Vaishnavi, et al.
Veröffentlicht: (2024)
von: Bhargava, Vaishnavi, et al.
Veröffentlicht: (2024)
Automating Code Adaptation for MLOps -- A Benchmarking Study on LLMs
von: Patel, Harsh, et al.
Veröffentlicht: (2024)
von: Patel, Harsh, et al.
Veröffentlicht: (2024)
MobileDev-Bench: A Benchmark for Issue Resolution in Mobile Application Development
von: Fakorede, Moshood A., et al.
Veröffentlicht: (2026)
von: Fakorede, Moshood A., et al.
Veröffentlicht: (2026)
Enabling Global, Human-Centered Explanations for LLMs:From Tokens to Interpretable Code and Test Generation
von: Khati, Dipin, et al.
Veröffentlicht: (2025)
von: Khati, Dipin, et al.
Veröffentlicht: (2025)
Think Anywhere in Code Generation
von: Jiang, Xue, et al.
Veröffentlicht: (2026)
von: Jiang, Xue, et al.
Veröffentlicht: (2026)
Operational Robustness of LLMs on Code Generation
von: Paul, Debalina Ghosh, et al.
Veröffentlicht: (2026)
von: Paul, Debalina Ghosh, et al.
Veröffentlicht: (2026)
Understanding Robustness of Model Editing in Code LLMs
von: Chhetri, Vinaik, et al.
Veröffentlicht: (2025)
von: Chhetri, Vinaik, et al.
Veröffentlicht: (2025)
Renaissance of Literate Programming in the Era of LLMs: Enhancing LLM-Based Code Generation in Large-Scale Projects
von: Zhang, Wuyang, et al.
Veröffentlicht: (2024)
von: Zhang, Wuyang, et al.
Veröffentlicht: (2024)
PromSec: Prompt Optimization for Secure Generation of Functional Source Code with Large Language Models (LLMs)
von: Nazzal, Mahmoud, et al.
Veröffentlicht: (2024)
von: Nazzal, Mahmoud, et al.
Veröffentlicht: (2024)
From Trace to Line: LLM Agent for Real-World OSS Vulnerability Localization
von: Xi, Haoran, et al.
Veröffentlicht: (2025)
von: Xi, Haoran, et al.
Veröffentlicht: (2025)
Unsupervised Evaluation of Code LLMs with Round-Trip Correctness
von: Allamanis, Miltiadis, et al.
Veröffentlicht: (2024)
von: Allamanis, Miltiadis, et al.
Veröffentlicht: (2024)
JavaBench: A Benchmark of Object-Oriented Code Generation for Evaluating Large Language Models
von: Cao, Jialun, et al.
Veröffentlicht: (2024)
von: Cao, Jialun, et al.
Veröffentlicht: (2024)
CodeTaste: Can LLMs Generate Human-Level Code Refactorings?
von: Thillen, Alex, et al.
Veröffentlicht: (2026)
von: Thillen, Alex, et al.
Veröffentlicht: (2026)
MermaidSeqBench: An Evaluation Benchmark for NL-to-Mermaid Sequence Diagram Generation
von: Shbita, Basel, et al.
Veröffentlicht: (2025)
von: Shbita, Basel, et al.
Veröffentlicht: (2025)
Mechanistic Interpretability of Code Correctness in LLMs via Sparse Autoencoders
von: Tahimic, Kriz, et al.
Veröffentlicht: (2025)
von: Tahimic, Kriz, et al.
Veröffentlicht: (2025)
SWE Atlas: Benchmarking Coding Agents Beyond Issue Resolution
von: Raghavendra, Mohit, et al.
Veröffentlicht: (2026)
von: Raghavendra, Mohit, et al.
Veröffentlicht: (2026)
SWE-Bench+: Enhanced Coding Benchmark for LLMs
von: Aleithan, Reem, et al.
Veröffentlicht: (2024)
von: Aleithan, Reem, et al.
Veröffentlicht: (2024)
TDD-Bench Verified: Can LLMs Generate Tests for Issues Before They Get Resolved?
von: Ahmed, Toufique, et al.
Veröffentlicht: (2024)
von: Ahmed, Toufique, et al.
Veröffentlicht: (2024)
Protocode: Prototype-Driven Interpretability for Code Generation in LLMs
von: Bodla, Krishna Vamshi, et al.
Veröffentlicht: (2025)
von: Bodla, Krishna Vamshi, et al.
Veröffentlicht: (2025)
Rethinking Repetition Problems of LLMs in Code Generation
von: Dong, Yihong, et al.
Veröffentlicht: (2025)
von: Dong, Yihong, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Enhanced Differential Testing in Emerging Database Systems
von: Jiang, Yuancheng, et al.
Veröffentlicht: (2025) -
AICD Bench: A Challenging Benchmark for AI-Generated Code Detection
von: Orel, Daniil, et al.
Veröffentlicht: (2026) -
ThrowBench: Benchmarking LLMs by Predicting Runtime Exceptions
von: Prenner, Julian Aron, et al.
Veröffentlicht: (2025) -
Deep-Bench: Deep Learning Benchmark Dataset for Code Generation
von: Daghighfarsoodeh, Alireza, et al.
Veröffentlicht: (2025) -
Where Do LLMs Still Struggle? An In-Depth Analysis of Code Generation Benchmarks
von: Sharifloo, Amir Molzam, et al.
Veröffentlicht: (2025)