Assessing LLM code generation quality through path planning tasks
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Wanyi, Su, Meng-Wen, Cummings, Mary L. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PyBench: Evaluating LLM Agent on various real-world coding tasks
by: Zhang, Yaolun, et al.
Published: (2024)
by: Zhang, Yaolun, et al.
Published: (2024)
Can LLMs plan paths in the real world?
by: Chen, Wanyi, et al.
Published: (2024)
by: Chen, Wanyi, et al.
Published: (2024)
An evaluation of LLM code generation capabilities through graded exercises
by: Jiménez, Álvaro Barbero
Published: (2024)
by: Jiménez, Álvaro Barbero
Published: (2024)
Canonical Intermediate Representation for LLM-based optimization problem formulation and code generation
by: Lyu, Zhongyuan, et al.
Published: (2026)
by: Lyu, Zhongyuan, et al.
Published: (2026)
NeSy is alive and well: A LLM-driven symbolic approach for better code comment data generation and classification
by: Akl, Hanna Abi
Published: (2024)
by: Akl, Hanna Abi
Published: (2024)
WIP: Assessing the Effectiveness of ChatGPT in Preparatory Testing Activities
by: Haldar, Susmita, et al.
Published: (2025)
by: Haldar, Susmita, et al.
Published: (2025)
Lifecycle-Aware code generation: Leveraging Software Engineering Phases in LLMs
by: Xing, Xing, et al.
Published: (2025)
by: Xing, Xing, et al.
Published: (2025)
Evaluating perturbation robustness of generative systems that use COBOL code inputs
by: Ackerman, Samuel, et al.
Published: (2025)
by: Ackerman, Samuel, et al.
Published: (2025)
LLM as a code generator in Agile Model Driven Development
by: Sadik, Ahmed R., et al.
Published: (2024)
by: Sadik, Ahmed R., et al.
Published: (2024)
REMODEL-LLM: Transforming C code to Java using LLMs
by: Gupta, Aryan, et al.
Published: (2025)
by: Gupta, Aryan, et al.
Published: (2025)
CodeXEmbed: A Generalist Embedding Model Family for Multiligual and Multi-task Code Retrieval
by: Liu, Ye, et al.
Published: (2024)
by: Liu, Ye, et al.
Published: (2024)
The Tyranny of Possibilities in the Design of Task-Oriented LLM Systems: A Scoping Survey
by: Dhamani, Dhruv, et al.
Published: (2023)
by: Dhamani, Dhruv, et al.
Published: (2023)
Evaluating AI-generated code for C++, Fortran, Go, Java, Julia, Matlab, Python, R, and Rust
by: Diehl, Patrick, et al.
Published: (2024)
by: Diehl, Patrick, et al.
Published: (2024)
PELLI: Framework to effectively integrate LLMs for quality software generation
by: Krebs, Rasmus, et al.
Published: (2026)
by: Krebs, Rasmus, et al.
Published: (2026)
GeoAnalystBench: A GeoAI benchmark for assessing large language models for spatial analysis workflow and code generation
by: Zhang, Qianheng, et al.
Published: (2025)
by: Zhang, Qianheng, et al.
Published: (2025)
Quality Assurance of LLM-generated Code: Addressing Non-Functional Quality Characteristics
by: Sun, Xin, et al.
Published: (2025)
by: Sun, Xin, et al.
Published: (2025)
Green My LLM: Studying the key factors affecting the energy consumption of code assistants
by: Coignion, Tristan, et al.
Published: (2024)
by: Coignion, Tristan, et al.
Published: (2024)
CoCo-Bench: A Comprehensive Code Benchmark For Multi-task Large Language Model Evaluation
by: Yin, Wenjing, et al.
Published: (2025)
by: Yin, Wenjing, et al.
Published: (2025)
Identifying Performance-Sensitive Configurations in Software Systems through Code Analysis with LLM Agents
by: Wang, Zehao, et al.
Published: (2024)
by: Wang, Zehao, et al.
Published: (2024)
LLM-Based Test Case Generation in DBMS through Monte Carlo Tree Search
by: Chen, Yujia, et al.
Published: (2026)
by: Chen, Yujia, et al.
Published: (2026)
Towards LLM-generated explanations for Component-based Knowledge Graph Question Answering Systems
by: Schiese, Dennis, et al.
Published: (2025)
by: Schiese, Dennis, et al.
Published: (2025)
Diffusion is a code repair operator and generator
by: Singh, Mukul, et al.
Published: (2025)
by: Singh, Mukul, et al.
Published: (2025)
Design choices made by LLM-based test generators prevent them from finding bugs
by: Mathews, Noble Saji, et al.
Published: (2024)
by: Mathews, Noble Saji, et al.
Published: (2024)
Patched MOA: optimizing inference for diverse software development tasks
by: Sharma, Asankhaya
Published: (2024)
by: Sharma, Asankhaya
Published: (2024)
Patched RTC: evaluating LLMs for diverse software development tasks
by: Sharma, Asankhaya
Published: (2024)
by: Sharma, Asankhaya
Published: (2024)
An Empirical Study on LLM-based Agents for Automated Bug Fixing
by: Meng, Xiangxin, et al.
Published: (2024)
by: Meng, Xiangxin, et al.
Published: (2024)
Uncovering Intention through LLM-Driven Code Snippet Description Generation
by: Nugroho, Yusuf Sulistyo, et al.
Published: (2025)
by: Nugroho, Yusuf Sulistyo, et al.
Published: (2025)
SmartMLOps Studio: Design of an LLM-Integrated IDE with Automated MLOps Pipelines for Model Development and Monitoring
by: Jin, Jiawei, et al.
Published: (2025)
by: Jin, Jiawei, et al.
Published: (2025)
LLM-Explorer: Towards Efficient and Affordable LLM-based Exploration for Mobile Apps
by: Zhao, Shanhui, et al.
Published: (2025)
by: Zhao, Shanhui, et al.
Published: (2025)
LLM-based agents for automating the enhancement of user story quality: An early report
by: Zhang, Zheying, et al.
Published: (2024)
by: Zhang, Zheying, et al.
Published: (2024)
Where Code Meets Natural Language: Taxonomy-Driven Information Flow Analysis for LLM-Integrated Applications
by: Xu, Zihao, et al.
Published: (2026)
by: Xu, Zihao, et al.
Published: (2026)
ComBench: A Repo-level Real-world Benchmark for Compilation Error Repair
by: Li, Jia, et al.
Published: (2026)
by: Li, Jia, et al.
Published: (2026)
Evolution without an Oracle: Driving Effective Evolution with LLM Judges
by: Zhao, Zhe, et al.
Published: (2025)
by: Zhao, Zhe, et al.
Published: (2025)
Enhancing LLM-Based Coding Tools through Native Integration of IDE-Derived Static Context
by: Li, Yichen, et al.
Published: (2024)
by: Li, Yichen, et al.
Published: (2024)
SolidCoder: Bridging the Mental-Reality Gap in LLM Code Generation through Concrete Execution
by: Lee, Woojin, et al.
Published: (2026)
by: Lee, Woojin, et al.
Published: (2026)
Vul-R2: A Reasoning LLM for Automated Vulnerability Repair
by: Wen, Xin-Cheng, et al.
Published: (2025)
by: Wen, Xin-Cheng, et al.
Published: (2025)
LLM assisted web application functional requirements generation: A case study of four popular LLMs over a Mess Management System
by: Gupta, Rashmi, et al.
Published: (2025)
by: Gupta, Rashmi, et al.
Published: (2025)
PropertyGPT: LLM-driven Formal Verification of Smart Contracts through Retrieval-Augmented Property Generation
by: Liu, Ye, et al.
Published: (2024)
by: Liu, Ye, et al.
Published: (2024)
Generating Energy-efficient code with LLMs
by: Cappendijk, Tom, et al.
Published: (2024)
by: Cappendijk, Tom, et al.
Published: (2024)
Security of LLM-generated Code: A Comparative Analysis
by: Morkonda, Srivathsan G, et al.
Published: (2026)
by: Morkonda, Srivathsan G, et al.
Published: (2026)
Similar Items
-
PyBench: Evaluating LLM Agent on various real-world coding tasks
by: Zhang, Yaolun, et al.
Published: (2024) -
Can LLMs plan paths in the real world?
by: Chen, Wanyi, et al.
Published: (2024) -
An evaluation of LLM code generation capabilities through graded exercises
by: Jiménez, Álvaro Barbero
Published: (2024) -
Canonical Intermediate Representation for LLM-based optimization problem formulation and code generation
by: Lyu, Zhongyuan, et al.
Published: (2026) -
NeSy is alive and well: A LLM-driven symbolic approach for better code comment data generation and classification
by: Akl, Hanna Abi
Published: (2024)