Operational Robustness of LLMs on Code Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Paul, Debalina Ghosh, Zhu, Hong, Bayley, Ian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Investigating The Smells of LLM Generated Code
von: Paul, Debalina Ghosh, et al.
Veröffentlicht: (2025)
von: Paul, Debalina Ghosh, et al.
Veröffentlicht: (2025)
Benchmarks and Metrics for Evaluations of Code Generation: A Critical Review
von: Paul, Debalina Ghosh, et al.
Veröffentlicht: (2024)
von: Paul, Debalina Ghosh, et al.
Veröffentlicht: (2024)
ScenEval: A Benchmark for Scenario-Based Evaluation of Code Generation
von: Paul, Debalina Ghosh, et al.
Veröffentlicht: (2024)
von: Paul, Debalina Ghosh, et al.
Veröffentlicht: (2024)
CodeTaste: Can LLMs Generate Human-Level Code Refactorings?
von: Thillen, Alex, et al.
Veröffentlicht: (2026)
von: Thillen, Alex, et al.
Veröffentlicht: (2026)
Protocode: Prototype-Driven Interpretability for Code Generation in LLMs
von: Bodla, Krishna Vamshi, et al.
Veröffentlicht: (2025)
von: Bodla, Krishna Vamshi, et al.
Veröffentlicht: (2025)
Drawing Pandas: A Benchmark for LLMs in Generating Plotting Code
von: Galimzyanov, Timur, et al.
Veröffentlicht: (2024)
von: Galimzyanov, Timur, et al.
Veröffentlicht: (2024)
LiCoEval: Evaluating LLMs on License Compliance in Code Generation
von: Xu, Weiwei, et al.
Veröffentlicht: (2024)
von: Xu, Weiwei, et al.
Veröffentlicht: (2024)
On LLMs' Internal Representation of Code Correctness
von: Ribeiro, Francisco, et al.
Veröffentlicht: (2025)
von: Ribeiro, Francisco, et al.
Veröffentlicht: (2025)
Prism: Dynamic and Flexible Benchmarking of LLMs Code Generation with Monte Carlo Tree Search
von: Majdinasab, Vahid, et al.
Veröffentlicht: (2025)
von: Majdinasab, Vahid, et al.
Veröffentlicht: (2025)
Evaluating the Use of LLMs for Documentation to Code Traceability
von: Alor, Ebube, et al.
Veröffentlicht: (2025)
von: Alor, Ebube, et al.
Veröffentlicht: (2025)
Semantic Voting: Execution-Grounded Consensus for LLM Code Generation
von: Jiang, Shan, et al.
Veröffentlicht: (2026)
von: Jiang, Shan, et al.
Veröffentlicht: (2026)
LLMs in Coding and their Impact on the Commercial Software Engineering Landscape
von: Belozerov, Vladislav, et al.
Veröffentlicht: (2025)
von: Belozerov, Vladislav, et al.
Veröffentlicht: (2025)
The Struggles of LLMs in Cross-lingual Code Clone Detection
von: Moumoula, Micheline Bénédicte, et al.
Veröffentlicht: (2024)
von: Moumoula, Micheline Bénédicte, et al.
Veröffentlicht: (2024)
Automating Code Adaptation for MLOps -- A Benchmarking Study on LLMs
von: Patel, Harsh, et al.
Veröffentlicht: (2024)
von: Patel, Harsh, et al.
Veröffentlicht: (2024)
How Robustly do LLMs Understand Execution Semantics?
von: Spiess, Claudio, et al.
Veröffentlicht: (2026)
von: Spiess, Claudio, et al.
Veröffentlicht: (2026)
Rethinking Repetition Problems of LLMs in Code Generation
von: Dong, Yihong, et al.
Veröffentlicht: (2025)
von: Dong, Yihong, et al.
Veröffentlicht: (2025)
An Empirical Evaluation of Locally Deployed LLMs for Bug Detection in Python Code
von: Vulićević, Jelena Ilić
Veröffentlicht: (2026)
von: Vulićević, Jelena Ilić
Veröffentlicht: (2026)
SnipGen: A Mining Repository Framework for Evaluating LLMs for Code
von: Rodriguez-Cardenas, Daniel, et al.
Veröffentlicht: (2025)
von: Rodriguez-Cardenas, Daniel, et al.
Veröffentlicht: (2025)
Free and Customizable Code Documentation with LLMs: A Fine-Tuning Approach
von: Chakrabarty, Sayak, et al.
Veröffentlicht: (2024)
von: Chakrabarty, Sayak, et al.
Veröffentlicht: (2024)
Towards More Trustworthy and Interpretable LLMs for Code through Syntax-Grounded Explanations
von: Palacio, David N., et al.
Veröffentlicht: (2024)
von: Palacio, David N., et al.
Veröffentlicht: (2024)
Can Coding Agents Be General Agents?
von: Ivanov, Maksim, et al.
Veröffentlicht: (2026)
von: Ivanov, Maksim, et al.
Veröffentlicht: (2026)
Optimizing AI-Assisted Code Generation
von: Torka, Simon, et al.
Veröffentlicht: (2024)
von: Torka, Simon, et al.
Veröffentlicht: (2024)
Keeping Code-Aware LLMs Fresh: Full Refresh, In-Context Deltas, and Incremental Fine-Tuning
von: Sharma, Pradeep Kumar, et al.
Veröffentlicht: (2025)
von: Sharma, Pradeep Kumar, et al.
Veröffentlicht: (2025)
Code Generation by Differential Test Time Scaling
von: He, Yifeng, et al.
Veröffentlicht: (2026)
von: He, Yifeng, et al.
Veröffentlicht: (2026)
Clover: Closed-Loop Verifiable Code Generation
von: Sun, Chuyue, et al.
Veröffentlicht: (2023)
von: Sun, Chuyue, et al.
Veröffentlicht: (2023)
Functional Overlap Reranking for Neural Code Generation
von: To, Hung Quoc, et al.
Veröffentlicht: (2023)
von: To, Hung Quoc, et al.
Veröffentlicht: (2023)
A Theoretical Analysis of Test-Driven Code Generation
von: Menet, Nicolas, et al.
Veröffentlicht: (2026)
von: Menet, Nicolas, et al.
Veröffentlicht: (2026)
CodeGeeX: A Pre-Trained Model for Code Generation with Multilingual Benchmarking on HumanEval-X
von: Zheng, Qinkai, et al.
Veröffentlicht: (2023)
von: Zheng, Qinkai, et al.
Veröffentlicht: (2023)
Enhancing LLM-Based Test Generation by Eliminating Covered Code
von: Xu, WeiZhe, et al.
Veröffentlicht: (2026)
von: Xu, WeiZhe, et al.
Veröffentlicht: (2026)
Exploring Pass-Rate Reward in Reinforcement Learning for Code Generation
von: Li, Xin-Ye, et al.
Veröffentlicht: (2026)
von: Li, Xin-Ye, et al.
Veröffentlicht: (2026)
Deep-Bench: Deep Learning Benchmark Dataset for Code Generation
von: Daghighfarsoodeh, Alireza, et al.
Veröffentlicht: (2025)
von: Daghighfarsoodeh, Alireza, et al.
Veröffentlicht: (2025)
Supersonic: Learning to Generate Source Code Optimizations in C/C++
von: Chen, Zimin, et al.
Veröffentlicht: (2023)
von: Chen, Zimin, et al.
Veröffentlicht: (2023)
Co-Located Tests, Better AI Code: How Test Syntax Structure Affects Foundation Model Code Generation
von: Jacopin, Éric
Veröffentlicht: (2026)
von: Jacopin, Éric
Veröffentlicht: (2026)
MASTEST: A LLM-Based Multi-Agent System For RESTful API Tests
von: Han, Xiaoke, et al.
Veröffentlicht: (2025)
von: Han, Xiaoke, et al.
Veröffentlicht: (2025)
An Initial Exploration of Contrastive Prompt Tuning to Generate Energy-Efficient Code
von: Weidmann, Sophie, et al.
Veröffentlicht: (2026)
von: Weidmann, Sophie, et al.
Veröffentlicht: (2026)
VeriScale: Adversarial Test-Suite Scaling for Verifiable Code Generation
von: Bai, Yifan, et al.
Veröffentlicht: (2026)
von: Bai, Yifan, et al.
Veröffentlicht: (2026)
VeriContest: A Competitive-Programming Benchmark for Verifiable Code Generation
von: Xie, Zichen, et al.
Veröffentlicht: (2026)
von: Xie, Zichen, et al.
Veröffentlicht: (2026)
Promise and Peril of Collaborative Code Generation Models: Balancing Effectiveness and Memorization
von: Chen, Zhi, et al.
Veröffentlicht: (2024)
von: Chen, Zhi, et al.
Veröffentlicht: (2024)
Can LLMs Generate Architectural Design Decisions? -An Exploratory Empirical study
von: Dhar, Rudra, et al.
Veröffentlicht: (2024)
von: Dhar, Rudra, et al.
Veröffentlicht: (2024)
What You See Is Not Always What You Get: Evaluating GPT's Comprehension of Source Code
von: Wen, Jiawen, et al.
Veröffentlicht: (2024)
von: Wen, Jiawen, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Investigating The Smells of LLM Generated Code
von: Paul, Debalina Ghosh, et al.
Veröffentlicht: (2025) -
Benchmarks and Metrics for Evaluations of Code Generation: A Critical Review
von: Paul, Debalina Ghosh, et al.
Veröffentlicht: (2024) -
ScenEval: A Benchmark for Scenario-Based Evaluation of Code Generation
von: Paul, Debalina Ghosh, et al.
Veröffentlicht: (2024) -
CodeTaste: Can LLMs Generate Human-Level Code Refactorings?
von: Thillen, Alex, et al.
Veröffentlicht: (2026) -
Protocode: Prototype-Driven Interpretability for Code Generation in LLMs
von: Bodla, Krishna Vamshi, et al.
Veröffentlicht: (2025)