The Counterfeit Conundrum: Can Code Language Models Grasp the Nuances of Their Incorrect Generations?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gu, Alex, Li, Wen-Ding, Jain, Naman, Olausson, Theo X., Lee, Celine, Sen, Koushik, Solar-Lezama, Armando |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
von: Jain, Naman, et al.
Veröffentlicht: (2024)
von: Jain, Naman, et al.
Veröffentlicht: (2024)
Is Self-Repair a Silver Bullet for Code Generation?
von: Olausson, Theo X., et al.
Veröffentlicht: (2023)
von: Olausson, Theo X., et al.
Veröffentlicht: (2023)
Challenges and Paths Towards AI for Software Engineering
von: Gu, Alex, et al.
Veröffentlicht: (2025)
von: Gu, Alex, et al.
Veröffentlicht: (2025)
CRUXEval: A Benchmark for Code Reasoning, Understanding and Execution
von: Gu, Alex, et al.
Veröffentlicht: (2024)
von: Gu, Alex, et al.
Veröffentlicht: (2024)
KISS Sorcar: A Stupidly-Simple General-Purpose and Software Engineering AI Assistant
von: Sen, Koushik
Veröffentlicht: (2026)
von: Sen, Koushik
Veröffentlicht: (2026)
R2E-Gym: Procedural Environments and Hybrid Verifiers for Scaling Open-Weights SWE Agents
von: Jain, Naman, et al.
Veröffentlicht: (2025)
von: Jain, Naman, et al.
Veröffentlicht: (2025)
Measuring the Influence of Incorrect Code on Test Generation
von: Huang, Dong, et al.
Veröffentlicht: (2024)
von: Huang, Dong, et al.
Veröffentlicht: (2024)
Syzygy: Dual Code-Test C to (safe) Rust Translation using LLMs and Dynamic Analysis
von: Shetty, Manish, et al.
Veröffentlicht: (2024)
von: Shetty, Manish, et al.
Veröffentlicht: (2024)
Code Copycat Conundrum: Demystifying Repetition in LLM-based Code Generation
von: Liu, Mingwei, et al.
Veröffentlicht: (2025)
von: Liu, Mingwei, et al.
Veröffentlicht: (2025)
GSO: Challenging Software Optimization Tasks for Evaluating SWE-Agents
von: Shetty, Manish, et al.
Veröffentlicht: (2025)
von: Shetty, Manish, et al.
Veröffentlicht: (2025)
Security Debt in Practice: Nuanced Insights from Practitioners
von: Boufaied, Chaima, et al.
Veröffentlicht: (2025)
von: Boufaied, Chaima, et al.
Veröffentlicht: (2025)
Copilot Arena: A Platform for Code LLM Evaluation in the Wild
von: Chi, Wayne, et al.
Veröffentlicht: (2025)
von: Chi, Wayne, et al.
Veröffentlicht: (2025)
SelfCodeAlign: Self-Alignment for Code Generation
von: Wei, Yuxiang, et al.
Veröffentlicht: (2024)
von: Wei, Yuxiang, et al.
Veröffentlicht: (2024)
LLM4Fuzz: Guided Fuzzing of Smart Contracts with Large Language Models
von: Shou, Chaofan, et al.
Veröffentlicht: (2024)
von: Shou, Chaofan, et al.
Veröffentlicht: (2024)
Agentic Much? Adoption of Coding Agents on GitHub
von: Robbes, Romain, et al.
Veröffentlicht: (2026)
von: Robbes, Romain, et al.
Veröffentlicht: (2026)
Promises, Perils, and (Timely) Heuristics for Mining Coding Agent Activity
von: Robbes, Romain, et al.
Veröffentlicht: (2026)
von: Robbes, Romain, et al.
Veröffentlicht: (2026)
ChatGPT Incorrectness Detection in Software Reviews
von: Tanzil, Minaoar Hossain, et al.
Veröffentlicht: (2024)
von: Tanzil, Minaoar Hossain, et al.
Veröffentlicht: (2024)
Can Large Language Models Generate Geospatial Code?
von: Hou, Shuyang, et al.
Veröffentlicht: (2024)
von: Hou, Shuyang, et al.
Veröffentlicht: (2024)
CodeTaste: Can LLMs Generate Human-Level Code Refactorings?
von: Thillen, Alex, et al.
Veröffentlicht: (2026)
von: Thillen, Alex, et al.
Veröffentlicht: (2026)
Can Large Language Models Serve as Evaluators for Code Summarization?
von: Wu, Yang, et al.
Veröffentlicht: (2024)
von: Wu, Yang, et al.
Veröffentlicht: (2024)
DeCon: Detecting Incorrect Assertions via Postconditions Generated by a Large Language Model
von: Yu, Hao, et al.
Veröffentlicht: (2025)
von: Yu, Hao, et al.
Veröffentlicht: (2025)
CodeScore: Evaluating Code Generation by Learning Code Execution
von: Dong, Yihong, et al.
Veröffentlicht: (2023)
von: Dong, Yihong, et al.
Veröffentlicht: (2023)
I Can't Share Code, but I need Translation -- An Empirical Study on Code Translation through Federated LLM
von: Kumar, Jahnavi, et al.
Veröffentlicht: (2025)
von: Kumar, Jahnavi, et al.
Veröffentlicht: (2025)
Can We Identify Stack Overflow Questions Requiring Code Snippets? Investigating the Cause & Effect of Missing Code Snippets
von: Mondal, Saikat, et al.
Veröffentlicht: (2024)
von: Mondal, Saikat, et al.
Veröffentlicht: (2024)
Can LLMs be Effective Code Contributors? A Study on Open-source Projects
von: Chong, Chun Jie, et al.
Veröffentlicht: (2026)
von: Chong, Chun Jie, et al.
Veröffentlicht: (2026)
RepoZero: Can LLMs Generate a Code Repository from Scratch?
von: Zhang, Zhaoxi, et al.
Veröffentlicht: (2026)
von: Zhang, Zhaoxi, et al.
Veröffentlicht: (2026)
Programming Language Confusion: When Code LLMs Can't Keep their Languages Straight
von: Moumoula, Micheline Bénédicte, et al.
Veröffentlicht: (2025)
von: Moumoula, Micheline Bénédicte, et al.
Veröffentlicht: (2025)
The Impact of the COVID-19 Pandemic on Women's Contribution to Public Code
von: Casanueva, Annalí, et al.
Veröffentlicht: (2024)
von: Casanueva, Annalí, et al.
Veröffentlicht: (2024)
Zero-Shot Code Representation Learning via Prompt Tuning
von: Cui, Nan, et al.
Veröffentlicht: (2024)
von: Cui, Nan, et al.
Veröffentlicht: (2024)
ZeroCoder: Can LLMs Improve Code Generation Without Ground-Truth Supervision?
von: Fan, Lishui, et al.
Veröffentlicht: (2026)
von: Fan, Lishui, et al.
Veröffentlicht: (2026)
SWE-Perf: Can Language Models Optimize Code Performance on Real-World Repositories?
von: He, Xinyi, et al.
Veröffentlicht: (2025)
von: He, Xinyi, et al.
Veröffentlicht: (2025)
Can LLMs Solve Science or Just Write Code? Evaluating Quantum Solver Generation
von: Baresi, Luciano, et al.
Veröffentlicht: (2026)
von: Baresi, Luciano, et al.
Veröffentlicht: (2026)
AVX / NEON Intrinsic Functions: When Should They Be Used?
von: Boivin, Théo, et al.
Veröffentlicht: (2026)
von: Boivin, Théo, et al.
Veröffentlicht: (2026)
CodeRAG-Bench: Can Retrieval Augment Code Generation?
von: Wang, Zora Zhiruo, et al.
Veröffentlicht: (2024)
von: Wang, Zora Zhiruo, et al.
Veröffentlicht: (2024)
Multiversion Hindsight Logging for Continuous Training
von: Garcia, Rolando, et al.
Veröffentlicht: (2023)
von: Garcia, Rolando, et al.
Veröffentlicht: (2023)
Multi-LLM Orchestration for High-Quality Code Generation: Exploiting Complementary Model Strengths
von: Chen, Huashan, et al.
Veröffentlicht: (2025)
von: Chen, Huashan, et al.
Veröffentlicht: (2025)
Cross-Domain Deep Code Search with Meta Learning
von: Chai, Yitian, et al.
Veröffentlicht: (2022)
von: Chai, Yitian, et al.
Veröffentlicht: (2022)
Which Code Statements Implement Privacy Behaviors in Android Applications?
von: Su, Chia-Yi, et al.
Veröffentlicht: (2025)
von: Su, Chia-Yi, et al.
Veröffentlicht: (2025)
A Multi-Language Perspective on the Robustness of LLM Code Generation
von: Rabbi, Fazle, et al.
Veröffentlicht: (2025)
von: Rabbi, Fazle, et al.
Veröffentlicht: (2025)
Can Emulating Semantic Translation Help LLMs with Code Translation? A Study Based on Pseudocode
von: Chen, Songqiang, et al.
Veröffentlicht: (2025)
von: Chen, Songqiang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
von: Jain, Naman, et al.
Veröffentlicht: (2024) -
Is Self-Repair a Silver Bullet for Code Generation?
von: Olausson, Theo X., et al.
Veröffentlicht: (2023) -
Challenges and Paths Towards AI for Software Engineering
von: Gu, Alex, et al.
Veröffentlicht: (2025) -
CRUXEval: A Benchmark for Code Reasoning, Understanding and Execution
von: Gu, Alex, et al.
Veröffentlicht: (2024) -
KISS Sorcar: A Stupidly-Simple General-Purpose and Software Engineering AI Assistant
von: Sen, Koushik
Veröffentlicht: (2026)