TaskEval: Assessing Difficulty of Code Generation Tasks for Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Tambon, Florian, Nikanjam, Amin, Zid, Cyrine, Khomh, Foutse, Antoniol, Giuliano |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Bugs in Large Language Models Generated Code: An Empirical Study
by: Tambon, Florian, et al.
Published: (2024)
by: Tambon, Florian, et al.
Published: (2024)
GIST: Generated Inputs Sets Transferability in Deep Learning
by: Tambon, Florian, et al.
Published: (2023)
by: Tambon, Florian, et al.
Published: (2023)
Prism: Dynamic and Flexible Benchmarking of LLMs Code Generation with Monte Carlo Tree Search
by: Majdinasab, Vahid, et al.
Published: (2025)
by: Majdinasab, Vahid, et al.
Published: (2025)
TaskEval: Synthesised Evaluation for Foundation-Model Tasks
by: Widanapathiranage, Dilani, et al.
Published: (2025)
by: Widanapathiranage, Dilani, et al.
Published: (2025)
Chain of Targeted Verification Questions to Improve the Reliability of Code Generated by LLMs
by: Ngassom, Sylvain Kouemo, et al.
Published: (2024)
by: Ngassom, Sylvain Kouemo, et al.
Published: (2024)
ReCatcher: Towards LLMs Regression Testing for Code Generation
by: Abbassi, Altaf Allah, et al.
Published: (2025)
by: Abbassi, Altaf Allah, et al.
Published: (2025)
DeepCodeProbe: Towards Understanding What Models Trained on Code Learn
by: Majdinasab, Vahid, et al.
Published: (2024)
by: Majdinasab, Vahid, et al.
Published: (2024)
Common Challenges of Deep Reinforcement Learning Applications Development: An Empirical Study
by: Morovati, Mohammad Mehdi, et al.
Published: (2023)
by: Morovati, Mohammad Mehdi, et al.
Published: (2023)
Trained Without My Consent: Detecting Code Inclusion In Language Models Trained on Code
by: Majdinasab, Vahid, et al.
Published: (2024)
by: Majdinasab, Vahid, et al.
Published: (2024)
Performance Smells in ML and Non-ML Python Projects: A Comparative Study
by: Belias, François, et al.
Published: (2025)
by: Belias, François, et al.
Published: (2025)
Inferring Code Correctness from Specification
by: Florian, Tambon, et al.
Published: (2026)
by: Florian, Tambon, et al.
Published: (2026)
A Taxonomy of Inefficiencies in LLM-Generated Python Code
by: Abbassi, Altaf Allah, et al.
Published: (2025)
by: Abbassi, Altaf Allah, et al.
Published: (2025)
A Survey of Bugs in AI-Generated Code
by: Gao, Ruofan, et al.
Published: (2025)
by: Gao, Ruofan, et al.
Published: (2025)
PathOCL: Path-Based Prompt Augmentation for OCL Generation with GPT-4
by: Abukhalaf, Seif, et al.
Published: (2024)
by: Abukhalaf, Seif, et al.
Published: (2024)
Fault Localization in Deep Learning-based Software: A System-level Approach
by: Morovati, Mohammad Mehdi, et al.
Published: (2024)
by: Morovati, Mohammad Mehdi, et al.
Published: (2024)
Toward Debugging Deep Reinforcement Learning Programs with RLExplorer
by: Bouchoucha, Rached, et al.
Published: (2024)
by: Bouchoucha, Rached, et al.
Published: (2024)
An Efficient Model Maintenance Approach for MLOps
by: Majidi, Forough, et al.
Published: (2024)
by: Majidi, Forough, et al.
Published: (2024)
Task Abstention for Large Language Models in Code Generation
by: Zhou, Yanke, et al.
Published: (2026)
by: Zhou, Yanke, et al.
Published: (2026)
RefAgent: A Multi-agent LLM-based Framework for Automatic Software Refactoring
by: Oueslati, Khouloud, et al.
Published: (2025)
by: Oueslati, Khouloud, et al.
Published: (2025)
Improving the Robustness of Large Language Models for Code Tasks via Fine-tuning with Perturbed Data
by: Liu, Yang, et al.
Published: (2026)
by: Liu, Yang, et al.
Published: (2026)
LLMs and Stack Overflow Discussions: Reliability, Impact, and Challenges
by: Da Silva, Leuson, et al.
Published: (2024)
by: Da Silva, Leuson, et al.
Published: (2024)
GeoCode-GPT: A Large Language Model for Geospatial Code Generation Tasks
by: Hou, Shuyang, et al.
Published: (2024)
by: Hou, Shuyang, et al.
Published: (2024)
Machine Learning Robustness: A Primer
by: Braiek, Houssem Ben, et al.
Published: (2024)
by: Braiek, Houssem Ben, et al.
Published: (2024)
SDLog: A Deep Learning Framework for Detecting Sensitive Information in Software Logs
by: Aghili, Roozbeh, et al.
Published: (2025)
by: Aghili, Roozbeh, et al.
Published: (2025)
Continuously Learning Bug Locations
by: Mindom, Paulina Stevia Nouwou, et al.
Published: (2024)
by: Mindom, Paulina Stevia Nouwou, et al.
Published: (2024)
What Information Contributes to Log-based Anomaly Detection? Insights from a Configurable Transformer-Based Approach
by: Wu, Xingfang, et al.
Published: (2024)
by: Wu, Xingfang, et al.
Published: (2024)
Deep Learning Model Reuse in the HuggingFace Community: Challenges, Benefit and Trends
by: Taraghi, Mina, et al.
Published: (2024)
by: Taraghi, Mina, et al.
Published: (2024)
ModiGen: A Large Language Model-Based Workflow for Multi-Task Modelica Code Generation
by: Xiang, Jiahui, et al.
Published: (2025)
by: Xiang, Jiahui, et al.
Published: (2025)
Adversarial Attack Classification and Robustness Testing for Large Language Models for Code
by: Liu, Yang, et al.
Published: (2025)
by: Liu, Yang, et al.
Published: (2025)
A Tertiary Review of Large Language Model-Based Code Generating Tasks: Trends, Challenges, and Future Directions
by: Chochlov, Muslim, et al.
Published: (2026)
by: Chochlov, Muslim, et al.
Published: (2026)
Code Red! On the Harmfulness of Applying Off-the-shelf Large Language Models to Programming Tasks
by: Al-Kaswan, Ali, et al.
Published: (2025)
by: Al-Kaswan, Ali, et al.
Published: (2025)
AdaptEval: A Benchmark for Evaluating Large Language Models on Code Snippet Adaptation
by: Zhang, Tanghaoran, et al.
Published: (2026)
by: Zhang, Tanghaoran, et al.
Published: (2026)
Imitation Game: Reproducing Deep Learning Bugs Leveraging an Intelligent Agent
by: Shah, Mehil B, et al.
Published: (2025)
by: Shah, Mehil B, et al.
Published: (2025)
Asm2SrcEval: Evaluating Large Language Models for Assembly-to-Source Code Translation
by: Hamedi, Parisa, et al.
Published: (2025)
by: Hamedi, Parisa, et al.
Published: (2025)
Assessing the Code Clone Detection Capability of Large Language Models
by: Zhang, Zixian, et al.
Published: (2024)
by: Zhang, Zixian, et al.
Published: (2024)
VHDL-Eval: A Framework for Evaluating Large Language Models in VHDL Code Generation
by: Vijayaraghavan, Prashanth, et al.
Published: (2024)
by: Vijayaraghavan, Prashanth, et al.
Published: (2024)
Automated Benchmark Generation for Repository-Level Coding Tasks
by: Vergopoulos, Konstantinos, et al.
Published: (2025)
by: Vergopoulos, Konstantinos, et al.
Published: (2025)
ConCodeEval: Evaluating Large Language Models for Code Constraints in Domain-Specific Languages
by: Kammakomati, Mehant, et al.
Published: (2024)
by: Kammakomati, Mehant, et al.
Published: (2024)
Advancing Language Models for Code-related Tasks
by: Tian, Zhao
Published: (2026)
by: Tian, Zhao
Published: (2026)
Benchmarking Large Language Models with Integer Sequence Generation Tasks
by: O'Malley, Daniel, et al.
Published: (2024)
by: O'Malley, Daniel, et al.
Published: (2024)
Similar Items
-
Bugs in Large Language Models Generated Code: An Empirical Study
by: Tambon, Florian, et al.
Published: (2024) -
GIST: Generated Inputs Sets Transferability in Deep Learning
by: Tambon, Florian, et al.
Published: (2023) -
Prism: Dynamic and Flexible Benchmarking of LLMs Code Generation with Monte Carlo Tree Search
by: Majdinasab, Vahid, et al.
Published: (2025) -
TaskEval: Synthesised Evaluation for Foundation-Model Tasks
by: Widanapathiranage, Dilani, et al.
Published: (2025) -
Chain of Targeted Verification Questions to Improve the Reliability of Code Generated by LLMs
by: Ngassom, Sylvain Kouemo, et al.
Published: (2024)