Human-Like Code Quality Evaluation through LLM-based Recursive Semantic Comprehension
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Fangzhou, Zhang, Sai, Xing, Zhenchang, Zhang, Xiaowang, Han, Yahong, Feng, Zhiyong |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A test-free semantic mistakes localization framework in Neural Code Translation
by: Chen, Lei, et al.
Published: (2024)
by: Chen, Lei, et al.
Published: (2024)
Empowering Agile-Based Generative Software Development through Human-AI Teamwork
by: Zhang, Sai, et al.
Published: (2024)
by: Zhang, Sai, et al.
Published: (2024)
Think Like an Engineer: A Neuro-Symbolic Collaboration Agent for Generative Software Requirements Elicitation and Self-Review
by: Zhang, Sai, et al.
Published: (2025)
by: Zhang, Sai, et al.
Published: (2025)
Do Chase Your Tail! Missing Key Aspects Augmentation in Textual Vulnerability Descriptions of Long-tail Software through Feature Inference
by: Han, Linyi, et al.
Published: (2024)
by: Han, Linyi, et al.
Published: (2024)
Domain-constrained Synthesis of Inconsistent Key Aspects in Textual Vulnerability Descriptions
by: Han, Linyi, et al.
Published: (2025)
by: Han, Linyi, et al.
Published: (2025)
Deployability-Centric Infrastructure-as-Code Generation: Fail, Learn, Refine, and Succeed through LLM-Empowered DevOps Simulation
by: Zhang, Tianyi, et al.
Published: (2025)
by: Zhang, Tianyi, et al.
Published: (2025)
Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation
by: Zhang, Binquan, et al.
Published: (2025)
by: Zhang, Binquan, et al.
Published: (2025)
Still Manual? Automated Linter Configuration via DSL-Based LLM Compilation of Coding Standards
by: Zhang, Zejun, et al.
Published: (2026)
by: Zhang, Zejun, et al.
Published: (2026)
A Large-scale Investigation of Semantically Incompatible APIs behind Compatibility Issues in Android Apps
by: Pan, Shidong, et al.
Published: (2024)
by: Pan, Shidong, et al.
Published: (2024)
Evaluation-Driven Development and Operations of LLM Agents: A Process Model and Reference Architecture
by: Xia, Boming, et al.
Published: (2024)
by: Xia, Boming, et al.
Published: (2024)
Re-Evaluating Code LLM Benchmarks Under Semantic Mutation
by: Pan, Zhiyuan, et al.
Published: (2025)
by: Pan, Zhiyuan, et al.
Published: (2025)
Refactoring to Pythonic Idioms: A Hybrid Knowledge-Driven Approach Leveraging Large Language Models
by: Zhang, Zejun, et al.
Published: (2024)
by: Zhang, Zejun, et al.
Published: (2024)
LLMs Meet Library Evolution: Evaluating Deprecated API Usage in LLM-based Code Completion
by: Wang, Chong, et al.
Published: (2024)
by: Wang, Chong, et al.
Published: (2024)
IntrinTrans: LLM-based Intrinsic Code Translator for RISC-V Vector
by: Han, Liutong, et al.
Published: (2025)
by: Han, Liutong, et al.
Published: (2025)
On the Quality of AI-Generated Source Code Comments: A Comprehensive Evaluation
by: Guelman, Ian, et al.
Published: (2024)
by: Guelman, Ian, et al.
Published: (2024)
Compiling Code LLMs into Lightweight Executables
by: Shi, Jieke, et al.
Published: (2026)
by: Shi, Jieke, et al.
Published: (2026)
Code Refactoring with LLM: A Comprehensive Evaluation With Few-Shot Settings
by: Tapader, Md. Raihan, et al.
Published: (2025)
by: Tapader, Md. Raihan, et al.
Published: (2025)
LogiAgent: Automated Logical Testing for REST Systems with LLM-Based Multi-Agents
by: Zhang, Ke, et al.
Published: (2025)
by: Zhang, Ke, et al.
Published: (2025)
Simplicity by Obfuscation: Evaluating LLM-Driven Code Transformation with Semantic Elasticity
by: De Tomasi, Lorenzo, et al.
Published: (2025)
by: De Tomasi, Lorenzo, et al.
Published: (2025)
A Comprehensive Evaluation of Parameter-Efficient Fine-Tuning on Code Smell Detection
by: Zhang, Beiqi, et al.
Published: (2024)
by: Zhang, Beiqi, et al.
Published: (2024)
CoRE: A Fine-Grained Code Reasoning Benchmark Beyond Output Prediction
by: Gao, Jun, et al.
Published: (2026)
by: Gao, Jun, et al.
Published: (2026)
A Benchmark for Localizing Code and Non-Code Issues in Software Projects
by: Zhang, Zejun, et al.
Published: (2025)
by: Zhang, Zejun, et al.
Published: (2025)
SeeAction: Towards Reverse Engineering How-What-Where of HCI Actions from Screencasts for UI Automation
by: Zhao, Dehai, et al.
Published: (2025)
by: Zhao, Dehai, et al.
Published: (2025)
SemGuard: Real-Time Semantic Evaluator for Correcting LLM-Generated Code
by: Wang, Qinglin, et al.
Published: (2025)
by: Wang, Qinglin, et al.
Published: (2025)
Benchmarking and Studying the LLM-based Code Review
by: Zeng, Zhengran, et al.
Published: (2025)
by: Zeng, Zhengran, et al.
Published: (2025)
A^3-CodGen: A Repository-Level Code Generation Framework for Code Reuse with Local-Aware, Global-Aware, and Third-Party-Library-Aware
by: Liao, Dianshu, et al.
Published: (2023)
by: Liao, Dianshu, et al.
Published: (2023)
Reducing Hallucinations in LLM-Generated Code via Semantic Triangulation
by: Dai, Yihan, et al.
Published: (2025)
by: Dai, Yihan, et al.
Published: (2025)
Debug Like a Human: Scaling LLM-based Fault Localization to Processor Design via Block-Level Instruction-Oriented Slicing
by: Liu, Zizhen, et al.
Published: (2026)
by: Liu, Zizhen, et al.
Published: (2026)
Enhancing Code Review through Fuzzing and Likely Invariants
by: Charoenwet, Wachiraphan, et al.
Published: (2025)
by: Charoenwet, Wachiraphan, et al.
Published: (2025)
From Code to Courtroom: LLMs as the New Software Judges
by: He, Junda, et al.
Published: (2025)
by: He, Junda, et al.
Published: (2025)
A Taxonomy of Foundation Model based Systems through the Lens of Software Architecture
by: Lu, Qinghua, et al.
Published: (2023)
by: Lu, Qinghua, et al.
Published: (2023)
The Effect of Code Obfuscation on Human Program Comprehension
by: Nguyen, Anh H. N., et al.
Published: (2026)
by: Nguyen, Anh H. N., et al.
Published: (2026)
Trust in Software Supply Chains: Blockchain-Enabled SBOM and the AIBOM Future
by: Xia, Boming, et al.
Published: (2023)
by: Xia, Boming, et al.
Published: (2023)
Decoding Human-LLM Collaboration in Coding: An Empirical Study of Multi-Turn Conversations in the Wild
by: Zhang, Binquan, et al.
Published: (2025)
by: Zhang, Binquan, et al.
Published: (2025)
Do AI Coding Agents Log Like Humans? An Empirical Study
by: Ouatiti, Youssef Esseddiq, et al.
Published: (2026)
by: Ouatiti, Youssef Esseddiq, et al.
Published: (2026)
A Survey of Code Review Benchmarks and Evaluation Practices in Pre-LLM and LLM Era
by: Khan, Taufiqul Islam, et al.
Published: (2026)
by: Khan, Taufiqul Islam, et al.
Published: (2026)
Automatically Benchmarking LLM Code Agents through Agent-Driven Annotation and Evaluation
by: Fu, Lingyue, et al.
Published: (2025)
by: Fu, Lingyue, et al.
Published: (2025)
LLM-as-a-Judge for Software Engineering: Literature Review, Vision, and the Road Ahead
by: He, Junda, et al.
Published: (2025)
by: He, Junda, et al.
Published: (2025)
TRACE: Evaluating Execution Efficiency of LLM-Based Code Translation
by: Gong, Zhihao, et al.
Published: (2026)
by: Gong, Zhihao, et al.
Published: (2026)
TRACE: Evaluating Execution Efficiency of LLM-Based Code Translation
by: Gong, Zhihao, et al.
Published: (2025)
by: Gong, Zhihao, et al.
Published: (2025)
Similar Items
-
A test-free semantic mistakes localization framework in Neural Code Translation
by: Chen, Lei, et al.
Published: (2024) -
Empowering Agile-Based Generative Software Development through Human-AI Teamwork
by: Zhang, Sai, et al.
Published: (2024) -
Think Like an Engineer: A Neuro-Symbolic Collaboration Agent for Generative Software Requirements Elicitation and Self-Review
by: Zhang, Sai, et al.
Published: (2025) -
Do Chase Your Tail! Missing Key Aspects Augmentation in Textual Vulnerability Descriptions of Long-tail Software through Feature Inference
by: Han, Linyi, et al.
Published: (2024) -
Domain-constrained Synthesis of Inconsistent Key Aspects in Textual Vulnerability Descriptions
by: Han, Linyi, et al.
Published: (2025)