Gespeichert in:
| Hauptverfasser: | Xu, Weiwei, Gao, Kai, He, Hao, Zhou, Minghui |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2408.02487 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CodeGeeX: A Pre-Trained Model for Code Generation with Multilingual Benchmarking on HumanEval-X
von: Zheng, Qinkai, et al.
Veröffentlicht: (2023)
von: Zheng, Qinkai, et al.
Veröffentlicht: (2023)
mcdok at SemEval-2026 Task 13: Finetuning LLMs for Detection of Machine-Generated Code
von: Skurla, Adam, et al.
Veröffentlicht: (2026)
von: Skurla, Adam, et al.
Veröffentlicht: (2026)
Evaluating the Use of LLMs for Documentation to Code Traceability
von: Alor, Ebube, et al.
Veröffentlicht: (2025)
von: Alor, Ebube, et al.
Veröffentlicht: (2025)
Operational Robustness of LLMs on Code Generation
von: Paul, Debalina Ghosh, et al.
Veröffentlicht: (2026)
von: Paul, Debalina Ghosh, et al.
Veröffentlicht: (2026)
The Readability Spectrum: Patterns, Issues, and Prompt Effects in LLM-Generated Code
von: Ye, Hengzhi, et al.
Veröffentlicht: (2026)
von: Ye, Hengzhi, et al.
Veröffentlicht: (2026)
Code Generation by Differential Test Time Scaling
von: He, Yifeng, et al.
Veröffentlicht: (2026)
von: He, Yifeng, et al.
Veröffentlicht: (2026)
CodeTaste: Can LLMs Generate Human-Level Code Refactorings?
von: Thillen, Alex, et al.
Veröffentlicht: (2026)
von: Thillen, Alex, et al.
Veröffentlicht: (2026)
Protocode: Prototype-Driven Interpretability for Code Generation in LLMs
von: Bodla, Krishna Vamshi, et al.
Veröffentlicht: (2025)
von: Bodla, Krishna Vamshi, et al.
Veröffentlicht: (2025)
An Empirical Evaluation of Locally Deployed LLMs for Bug Detection in Python Code
von: Vulićević, Jelena Ilić
Veröffentlicht: (2026)
von: Vulićević, Jelena Ilić
Veröffentlicht: (2026)
SnipGen: A Mining Repository Framework for Evaluating LLMs for Code
von: Rodriguez-Cardenas, Daniel, et al.
Veröffentlicht: (2025)
von: Rodriguez-Cardenas, Daniel, et al.
Veröffentlicht: (2025)
Drawing Pandas: A Benchmark for LLMs in Generating Plotting Code
von: Galimzyanov, Timur, et al.
Veröffentlicht: (2024)
von: Galimzyanov, Timur, et al.
Veröffentlicht: (2024)
GeoSQL-Eval: First Evaluation of LLMs on PostGIS-Based NL2GeoSQL Queries
von: Hou, Shuyang, et al.
Veröffentlicht: (2025)
von: Hou, Shuyang, et al.
Veröffentlicht: (2025)
A Comprehensive Framework for Evaluating API-oriented Code Generation in Large Language Models
von: Wu, Yixi, et al.
Veröffentlicht: (2024)
von: Wu, Yixi, et al.
Veröffentlicht: (2024)
On Randomness in Agentic Evals
von: Bjarnason, Bjarni Haukur, et al.
Veröffentlicht: (2026)
von: Bjarnason, Bjarni Haukur, et al.
Veröffentlicht: (2026)
CodeVisionary: An Agent-based Framework for Evaluating Large Language Models in Code Generation
von: Wang, Xinchen, et al.
Veröffentlicht: (2025)
von: Wang, Xinchen, et al.
Veröffentlicht: (2025)
Co-Located Tests, Better AI Code: How Test Syntax Structure Affects Foundation Model Code Generation
von: Jacopin, Éric
Veröffentlicht: (2026)
von: Jacopin, Éric
Veröffentlicht: (2026)
On LLMs' Internal Representation of Code Correctness
von: Ribeiro, Francisco, et al.
Veröffentlicht: (2025)
von: Ribeiro, Francisco, et al.
Veröffentlicht: (2025)
Prism: Dynamic and Flexible Benchmarking of LLMs Code Generation with Monte Carlo Tree Search
von: Majdinasab, Vahid, et al.
Veröffentlicht: (2025)
von: Majdinasab, Vahid, et al.
Veröffentlicht: (2025)
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs
von: Zhou, Zenghui, et al.
Veröffentlicht: (2026)
von: Zhou, Zenghui, et al.
Veröffentlicht: (2026)
The Compliance Paradox: Semantic-Instruction Decoupling in Automated Academic Code Evaluation
von: Sahoo, Devanshu, et al.
Veröffentlicht: (2026)
von: Sahoo, Devanshu, et al.
Veröffentlicht: (2026)
The Struggles of LLMs in Cross-lingual Code Clone Detection
von: Moumoula, Micheline Bénédicte, et al.
Veröffentlicht: (2024)
von: Moumoula, Micheline Bénédicte, et al.
Veröffentlicht: (2024)
LLMs in Coding and their Impact on the Commercial Software Engineering Landscape
von: Belozerov, Vladislav, et al.
Veröffentlicht: (2025)
von: Belozerov, Vladislav, et al.
Veröffentlicht: (2025)
Rethinking Repetition Problems of LLMs in Code Generation
von: Dong, Yihong, et al.
Veröffentlicht: (2025)
von: Dong, Yihong, et al.
Veröffentlicht: (2025)
Enhancing LLM-Based Test Generation by Eliminating Covered Code
von: Xu, WeiZhe, et al.
Veröffentlicht: (2026)
von: Xu, WeiZhe, et al.
Veröffentlicht: (2026)
LicenseGPT: A Fine-tuned Foundation Model for Publicly Available Dataset License Compliance
von: Tan, Jingwen, et al.
Veröffentlicht: (2024)
von: Tan, Jingwen, et al.
Veröffentlicht: (2024)
Automating Code Adaptation for MLOps -- A Benchmarking Study on LLMs
von: Patel, Harsh, et al.
Veröffentlicht: (2024)
von: Patel, Harsh, et al.
Veröffentlicht: (2024)
Evaluation of LLMs on Syntax-Aware Code Fill-in-the-Middle Tasks
von: Gong, Linyuan, et al.
Veröffentlicht: (2024)
von: Gong, Linyuan, et al.
Veröffentlicht: (2024)
Permissive-Washing in the Open AI Supply Chain: A Large-Scale Audit of License Integrity
von: Jewitt, James, et al.
Veröffentlicht: (2026)
von: Jewitt, James, et al.
Veröffentlicht: (2026)
Beyond Synthetic Benchmarks: Evaluating LLM Performance on Real-World Class-Level Code Generation
von: Rahman, Musfiqur, et al.
Veröffentlicht: (2025)
von: Rahman, Musfiqur, et al.
Veröffentlicht: (2025)
How Do Your Code LLMs Perform? Empowering Code Instruction Tuning with High-Quality Data
von: Wang, Yejie, et al.
Veröffentlicht: (2024)
von: Wang, Yejie, et al.
Veröffentlicht: (2024)
Free and Customizable Code Documentation with LLMs: A Fine-Tuning Approach
von: Chakrabarty, Sayak, et al.
Veröffentlicht: (2024)
von: Chakrabarty, Sayak, et al.
Veröffentlicht: (2024)
Towards More Trustworthy and Interpretable LLMs for Code through Syntax-Grounded Explanations
von: Palacio, David N., et al.
Veröffentlicht: (2024)
von: Palacio, David N., et al.
Veröffentlicht: (2024)
NoFunEval: Funny How Code LMs Falter on Requirements Beyond Functional Correctness
von: Singhal, Manav, et al.
Veröffentlicht: (2024)
von: Singhal, Manav, et al.
Veröffentlicht: (2024)
Optimizing AI-Assisted Code Generation
von: Torka, Simon, et al.
Veröffentlicht: (2024)
von: Torka, Simon, et al.
Veröffentlicht: (2024)
Can Coding Agents Be General Agents?
von: Ivanov, Maksim, et al.
Veröffentlicht: (2026)
von: Ivanov, Maksim, et al.
Veröffentlicht: (2026)
More with Less: An Empirical Study of Turn-Control Strategies for Efficient Coding Agents
von: Gao, Pengfei, et al.
Veröffentlicht: (2025)
von: Gao, Pengfei, et al.
Veröffentlicht: (2025)
Keeping Code-Aware LLMs Fresh: Full Refresh, In-Context Deltas, and Incremental Fine-Tuning
von: Sharma, Pradeep Kumar, et al.
Veröffentlicht: (2025)
von: Sharma, Pradeep Kumar, et al.
Veröffentlicht: (2025)
Clover: Closed-Loop Verifiable Code Generation
von: Sun, Chuyue, et al.
Veröffentlicht: (2023)
von: Sun, Chuyue, et al.
Veröffentlicht: (2023)
Functional Overlap Reranking for Neural Code Generation
von: To, Hung Quoc, et al.
Veröffentlicht: (2023)
von: To, Hung Quoc, et al.
Veröffentlicht: (2023)
LiveCodeBench Pro: How Do Olympiad Medalists Judge LLMs in Competitive Programming?
von: Zheng, Zihan, et al.
Veröffentlicht: (2025)
von: Zheng, Zihan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
CodeGeeX: A Pre-Trained Model for Code Generation with Multilingual Benchmarking on HumanEval-X
von: Zheng, Qinkai, et al.
Veröffentlicht: (2023) -
mcdok at SemEval-2026 Task 13: Finetuning LLMs for Detection of Machine-Generated Code
von: Skurla, Adam, et al.
Veröffentlicht: (2026) -
Evaluating the Use of LLMs for Documentation to Code Traceability
von: Alor, Ebube, et al.
Veröffentlicht: (2025) -
Operational Robustness of LLMs on Code Generation
von: Paul, Debalina Ghosh, et al.
Veröffentlicht: (2026) -
The Readability Spectrum: Patterns, Issues, and Prompt Effects in LLM-Generated Code
von: Ye, Hengzhi, et al.
Veröffentlicht: (2026)