Comparing Human and LLM Generated Code: The Jury is Still Out!
Fuente:
arXiv
Saved in:
| Main Authors: | Licorish, Sherlock A., Bajpai, Ansh, Arora, Chetan, Wang, Fanyu, Tantithamthavorn, Kla |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GBM Returns the Best Prediction Performance among Regression Approaches: A Case Study of Stack Overflow Code Quality
by: Licorish, Sherlock A., et al.
Published: (2025)
by: Licorish, Sherlock A., et al.
Published: (2025)
Assessing Reliability of Statistical Maximum Coverage Estimators in Fuzzing
by: Liyanage, Danushka, et al.
Published: (2025)
by: Liyanage, Danushka, et al.
Published: (2025)
A semantic mutation metric for metamorphic relation adequacy in scientific computing programs
by: Li, Meng, et al.
Published: (2026)
by: Li, Meng, et al.
Published: (2026)
Framework Matters: Energy Efficiency of UI Automation Testing Frameworks
by: Lagermann, Timmie M. R., et al.
Published: (2025)
by: Lagermann, Timmie M. R., et al.
Published: (2025)
Predictive Analytics for Collaborators Answers, Code Quality, and Dropout on Stack Overflow
by: Zolduoarrati, Elijah, et al.
Published: (2025)
by: Zolduoarrati, Elijah, et al.
Published: (2025)
Validating Formal Specifications with LLM-generated Test Cases
by: Cunha, Alcino, et al.
Published: (2025)
by: Cunha, Alcino, et al.
Published: (2025)
Testing SSD Firmware with State Data-Aware Fuzzing: Accelerating Coverage in Nondeterministic I/O Environments
by: Yoon, Gangho, et al.
Published: (2025)
by: Yoon, Gangho, et al.
Published: (2025)
SWE-ABS: Adversarial Benchmark Strengthening Exposes Inflated Success Rates on Test-based Benchmark
by: Yu, Boxi, et al.
Published: (2026)
by: Yu, Boxi, et al.
Published: (2026)
Combined Program Analysis Techniques: A Systematic Mapping Study
by: Braione, Pietro, et al.
Published: (2026)
by: Braione, Pietro, et al.
Published: (2026)
Automatically Detecting Numerical Instability in Machine Learning Applications via Soft Assertions
by: Sharmin, Shaila, et al.
Published: (2025)
by: Sharmin, Shaila, et al.
Published: (2025)
Evaluating the Overhead of the Performance Profiler Cloudprofiler With MooBench
by: Yang, Shinhyung, et al.
Published: (2024)
by: Yang, Shinhyung, et al.
Published: (2024)
Analyzing the Adoption of Database Management Systems Throughout the History of Open Source Projects
by: Paiva, Camila A., et al.
Published: (2026)
by: Paiva, Camila A., et al.
Published: (2026)
The Specification as Quality Gate: Three Hypotheses on AI-Assisted Code Review
by: Zietsman, Christo
Published: (2026)
by: Zietsman, Christo
Published: (2026)
Resilient Microservices: A Systematic Review of Recovery Patterns, Strategies, and Evaluation Frameworks
by: Mohammad, Muzeeb
Published: (2025)
by: Mohammad, Muzeeb
Published: (2025)
Synthesizing Test Cases for Narrowing Specification Candidates
by: Cunha, Alcino, et al.
Published: (2025)
by: Cunha, Alcino, et al.
Published: (2025)
The ACPATH Metric: Precise Estimation of the Number of Acyclic Paths in C-like Languages
by: Bagnara, Roberto, et al.
Published: (2016)
by: Bagnara, Roberto, et al.
Published: (2016)
Software Testing at the Network Layer: Automated HTTP API Quality Assessment and Security Analysis of Production Web Applications
by: Mughal, Ali Hassaan, et al.
Published: (2026)
by: Mughal, Ali Hassaan, et al.
Published: (2026)
LLMLOOP: Improving LLM-Generated Code and Tests through Automated Iterative Feedback Loops
by: Ravi, Ravin, et al.
Published: (2026)
by: Ravi, Ravin, et al.
Published: (2026)
Understanding the Effect of Agile Practice Quality on Software Product Quality
by: Licorish, Sherlock Anthony
Published: (2024)
by: Licorish, Sherlock Anthony
Published: (2024)
From Monolith to Microservices: A Comparative Evaluation of Decomposition Frameworks
by: Weerasinghe, Mineth, et al.
Published: (2026)
by: Weerasinghe, Mineth, et al.
Published: (2026)
Geographic Variation in Stack Overflow Code Quality: Evidence from a Cross-Regional Study of Coding Practices
by: Zolduoarrati, Elijah, et al.
Published: (2026)
by: Zolduoarrati, Elijah, et al.
Published: (2026)
Evaluating Cryptographic API Misuse Detectors for Go
by: Andersson, Vivi, et al.
Published: (2026)
by: Andersson, Vivi, et al.
Published: (2026)
Tractable Verification of Model Transformations: A Cutoff-Theorem Approach for DSLTrans
by: Lucio, Levi
Published: (2026)
by: Lucio, Levi
Published: (2026)
Feedback-Normalized Developer Memory for Reinforcement-Learning Coding Agents: A Safety-Gated MCP Architecture
by: Iscan, Mehmet
Published: (2026)
by: Iscan, Mehmet
Published: (2026)
Trace Validation of Unmodified Concurrent Systems with OmniLink
by: Hackett, Finn, et al.
Published: (2026)
by: Hackett, Finn, et al.
Published: (2026)
AIRA: AI-Induced Risk Audit: A Structured Inspection Framework for AI-Generated Code
by: Parris, William M.
Published: (2026)
by: Parris, William M.
Published: (2026)
InterEvo-TR: Interactive Evolutionary Test Generation With Readability Assessment
by: Delgado-Pérez, Pedro, et al.
Published: (2024)
by: Delgado-Pérez, Pedro, et al.
Published: (2024)
Towards a Probabilistic Framework for Analyzing and Improving LLM-Enabled Software
by: Baldonado, Juan Manuel, et al.
Published: (2025)
by: Baldonado, Juan Manuel, et al.
Published: (2025)
Context Engineering for Multi-Agent LLM Code Assistants Using Elicit, NotebookLM, ChatGPT, and Claude Code
by: Haseeb, Muhammad
Published: (2025)
by: Haseeb, Muhammad
Published: (2025)
Validating Solidity Code Defects using Symbolic and Concrete Execution powered by Large Language Models
by: Susan, Ştefan-Claudiu, et al.
Published: (2025)
by: Susan, Ştefan-Claudiu, et al.
Published: (2025)
Path-optimal symbolic execution of heap-manipulating programs
by: Braione, Pietro, et al.
Published: (2024)
by: Braione, Pietro, et al.
Published: (2024)
Stabilization Without Simplification: A Two-Dimensional Model of Software Evolution
by: Furukawa, Masaru
Published: (2026)
by: Furukawa, Masaru
Published: (2026)
Iterative Audit Convergence in LLM-Managed Multi-Agent Systems: A Case Study in Prompt Engineering Quality Assurance
by: Calboreanu, Elias
Published: (2026)
by: Calboreanu, Elias
Published: (2026)
GitHub Copilot and Developer Productivity: An Observational Dose-Response Analysis
by: Heilman, Alex, et al.
Published: (2026)
by: Heilman, Alex, et al.
Published: (2026)
On the Soundness and Consistency of LLM Agents for Executing Test Cases Written in Natural Language
by: Salva, Sébastien, et al.
Published: (2025)
by: Salva, Sébastien, et al.
Published: (2025)
Understanding and Reusing Test Suites Across Database Systems
by: Zhong, Suyang, et al.
Published: (2024)
by: Zhong, Suyang, et al.
Published: (2024)
BACE: LLM-based Code Generation through Bayesian Anchored Co-Evolution of Code and Test Populations
by: Silva, Kaushitha, et al.
Published: (2026)
by: Silva, Kaushitha, et al.
Published: (2026)
Making Software Metrics Useful
by: Tempero, Ewan, et al.
Published: (2026)
by: Tempero, Ewan, et al.
Published: (2026)
Automated Vulnerability Detection Using Deep Learning Technique
by: Yang, Guan-Yan, et al.
Published: (2024)
by: Yang, Guan-Yan, et al.
Published: (2024)
Monitoring Agentic Systems Before They're Reliable
by: Boston, Marisa Ferrara, et al.
Published: (2026)
by: Boston, Marisa Ferrara, et al.
Published: (2026)
Similar Items
-
GBM Returns the Best Prediction Performance among Regression Approaches: A Case Study of Stack Overflow Code Quality
by: Licorish, Sherlock A., et al.
Published: (2025) -
Assessing Reliability of Statistical Maximum Coverage Estimators in Fuzzing
by: Liyanage, Danushka, et al.
Published: (2025) -
A semantic mutation metric for metamorphic relation adequacy in scientific computing programs
by: Li, Meng, et al.
Published: (2026) -
Framework Matters: Energy Efficiency of UI Automation Testing Frameworks
by: Lagermann, Timmie M. R., et al.
Published: (2025) -
Predictive Analytics for Collaborators Answers, Code Quality, and Dropout on Stack Overflow
by: Zolduoarrati, Elijah, et al.
Published: (2025)