Identifying Inaccurate Descriptions in LLM-generated Code Comments via Test Execution
Fuente:
arXiv
Saved in:
| Main Authors: | Kang, Sungmin, Milliken, Louis, Yoo, Shin |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Beyond pip install: Evaluating LLM Agents for the Automated Installation of Python Projects
by: Milliken, Louis, et al.
Published: (2024)
by: Milliken, Louis, et al.
Published: (2024)
A Quantitative and Qualitative Evaluation of LLM-Based Explainable Fault Localization
by: Kang, Sungmin, et al.
Published: (2023)
by: Kang, Sungmin, et al.
Published: (2023)
Lachesis: Predicting LLM Inference Accuracy using Structural Properties of Reasoning Paths
by: Kim, Naryeong, et al.
Published: (2024)
by: Kim, Naryeong, et al.
Published: (2024)
COSMosFL: Ensemble of Small Language Models for Fault Localisation
by: Cho, Hyunjoon, et al.
Published: (2025)
by: Cho, Hyunjoon, et al.
Published: (2025)
Predictive Prompt Analysis
by: Lee, Jae Yong, et al.
Published: (2025)
by: Lee, Jae Yong, et al.
Published: (2025)
Finding the Needle in the Crash Stack: Industrial-Scale Crash Root Cause Localization with AutoCrashFL
by: Kang, Sungmin, et al.
Published: (2025)
by: Kang, Sungmin, et al.
Published: (2025)
Identifying Bug Inducing Commits by Combining Fault Localisation and Code Change Histories
by: An, Gabin, et al.
Published: (2025)
by: An, Gabin, et al.
Published: (2025)
Adaptive Testing for LLM-Based Applications: A Diversity-based Approach
by: Yoon, Juyeon, et al.
Published: (2025)
by: Yoon, Juyeon, et al.
Published: (2025)
AutoCodeSherpa: Symbolic Explanations in AI Coding Agents
by: Kang, Sungmin, et al.
Published: (2025)
by: Kang, Sungmin, et al.
Published: (2025)
METAMON: Finding Inconsistencies between Program Documentation and Behavior using Metamorphic LLM Queries
by: Lee, Hyeonseok, et al.
Published: (2025)
by: Lee, Hyeonseok, et al.
Published: (2025)
TRACE: Evaluating Execution Efficiency of LLM-Based Code Translation
by: Gong, Zhihao, et al.
Published: (2026)
by: Gong, Zhihao, et al.
Published: (2026)
TRACE: Evaluating Execution Efficiency of LLM-Based Code Translation
by: Gong, Zhihao, et al.
Published: (2025)
by: Gong, Zhihao, et al.
Published: (2025)
CSA-Trans: Code Structure Aware Transformer for AST
by: Oh, Saeyoon, et al.
Published: (2024)
by: Oh, Saeyoon, et al.
Published: (2024)
Clotho: Measuring Task-Specific Pre-Generation Test Adequacy for LLM Inputs
by: Yoon, Juyeon, et al.
Published: (2025)
by: Yoon, Juyeon, et al.
Published: (2025)
Testing Refactoring Engine via Historical Bug Report driven LLM
by: Wang, Haibo, et al.
Published: (2025)
by: Wang, Haibo, et al.
Published: (2025)
Demystifying Errors in LLM Reasoning Traces: An Empirical Study of Code Execution Simulation
by: Abdollahi, Mohammad, et al.
Published: (2025)
by: Abdollahi, Mohammad, et al.
Published: (2025)
HyClone: Bridging LLM Understanding and Dynamic Execution for Semantic Code Clone Detection
by: Liang, Yunhao, et al.
Published: (2025)
by: Liang, Yunhao, et al.
Published: (2025)
CodeScore: Evaluating Code Generation by Learning Code Execution
by: Dong, Yihong, et al.
Published: (2023)
by: Dong, Yihong, et al.
Published: (2023)
Automated Repair of Ambiguous Problem Descriptions for LLM-Based Code Generation
by: Jia, Haoxiang, et al.
Published: (2025)
by: Jia, Haoxiang, et al.
Published: (2025)
Impact of Comments on LLM Comprehension of Legacy Code
by: Sabetto, Rock, et al.
Published: (2025)
by: Sabetto, Rock, et al.
Published: (2025)
Structured Safety Auditing for Balancing Code Correctness and Content Safety in LLM-Generated Code
by: Tan, Honghao, et al.
Published: (2026)
by: Tan, Honghao, et al.
Published: (2026)
Revisiting "Revisiting Neuron Coverage for DNN Testing: A Layer-Wise and Distribution-Aware Criterion": A Critical Review and Implications on DNN Coverage Testing
by: Kim, Jinhan, et al.
Published: (2026)
by: Kim, Jinhan, et al.
Published: (2026)
Understanding on the Edge: LLM-generated Boundary Test Explanations
by: Akbarova, Sabinakhon, et al.
Published: (2026)
by: Akbarova, Sabinakhon, et al.
Published: (2026)
Test Wars: A Comparative Study of SBST, Symbolic Execution, and LLM-Based Approaches to Unit Test Generation
by: Abdullin, Azat, et al.
Published: (2025)
by: Abdullin, Azat, et al.
Published: (2025)
Python Symbolic Execution with LLM-powered Code Generation
by: Wang, Wenhan, et al.
Published: (2024)
by: Wang, Wenhan, et al.
Published: (2024)
Atropos: Improving Cost-Benefit Trade-off of LLM-based Agents under Self-Consistency with Early Termination and Model Hotswap
by: Kim, Naryeong, et al.
Published: (2026)
by: Kim, Naryeong, et al.
Published: (2026)
Detecting LLM-generated Code with Subtle Modification by Adversarial Training
by: Yin, Xin, et al.
Published: (2025)
by: Yin, Xin, et al.
Published: (2025)
GenX: Mastering Code and Test Generation with Execution Feedback
by: Wang, Nan, et al.
Published: (2024)
by: Wang, Nan, et al.
Published: (2024)
MuFF: Stable and Sensitive Post-training Mutation Testing for Deep Learning
by: Kim, Jinhan, et al.
Published: (2025)
by: Kim, Jinhan, et al.
Published: (2025)
Compiling Code LLMs into Lightweight Executables
by: Shi, Jieke, et al.
Published: (2026)
by: Shi, Jieke, et al.
Published: (2026)
Automated Harmfulness Testing for Code Large Language Models
by: Tan, Honghao, et al.
Published: (2025)
by: Tan, Honghao, et al.
Published: (2025)
ExeCoder: Empowering Large Language Models with Executability Representation for Code Translation
by: He, Minghua, et al.
Published: (2025)
by: He, Minghua, et al.
Published: (2025)
ChangeGuard: Validating Code Changes via Pairwise Learning-Guided Execution
by: Gröninger, Lars, et al.
Published: (2024)
by: Gröninger, Lars, et al.
Published: (2024)
Similar Pattern Annotation via Retrieval Knowledge for LLM-Based Test Code Fault Localization
by: Gharachorlu, Golnaz, et al.
Published: (2026)
by: Gharachorlu, Golnaz, et al.
Published: (2026)
Comment Traps: How Defective Commented-out Code Augment Defects in AI-Assisted Code Generation
by: Huang, Yuan, et al.
Published: (2025)
by: Huang, Yuan, et al.
Published: (2025)
Identifying Process Improvement Opportunities through Process Execution Benchmarking
by: Abb, Luka, et al.
Published: (2025)
by: Abb, Luka, et al.
Published: (2025)
Distinguishing LLM-generated from Human-written Code by Contrastive Learning
by: Xu, Xiaodan, et al.
Published: (2024)
by: Xu, Xiaodan, et al.
Published: (2024)
TransAgent: Enhancing LLM-Based Code Translation via Fine-Grained Execution Alignment
by: Yuan, Zhiqiang, et al.
Published: (2024)
by: Yuan, Zhiqiang, et al.
Published: (2024)
Scaling Mobile Chaos Testing with AI-Driven Test Execution
by: Marcano, Juan, et al.
Published: (2026)
by: Marcano, Juan, et al.
Published: (2026)
LLM-Guided Issue Generation from Uncovered Code Segments
by: Pressato, Diany, et al.
Published: (2026)
by: Pressato, Diany, et al.
Published: (2026)
Similar Items
-
Beyond pip install: Evaluating LLM Agents for the Automated Installation of Python Projects
by: Milliken, Louis, et al.
Published: (2024) -
A Quantitative and Qualitative Evaluation of LLM-Based Explainable Fault Localization
by: Kang, Sungmin, et al.
Published: (2023) -
Lachesis: Predicting LLM Inference Accuracy using Structural Properties of Reasoning Paths
by: Kim, Naryeong, et al.
Published: (2024) -
COSMosFL: Ensemble of Small Language Models for Fault Localisation
by: Cho, Hyunjoon, et al.
Published: (2025) -
Predictive Prompt Analysis
by: Lee, Jae Yong, et al.
Published: (2025)