Prism: Dynamic and Flexible Benchmarking of LLMs Code Generation with Monte Carlo Tree Search
Fuente:
arXiv
Salvato in:
| Autori principali: | Majdinasab, Vahid, Nikanjam, Amin, Khomh, Foutse |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
DeepCodeProbe: Towards Understanding What Models Trained on Code Learn
di: Majdinasab, Vahid, et al.
Pubblicazione: (2024)
di: Majdinasab, Vahid, et al.
Pubblicazione: (2024)
Trained Without My Consent: Detecting Code Inclusion In Language Models Trained on Code
di: Majdinasab, Vahid, et al.
Pubblicazione: (2024)
di: Majdinasab, Vahid, et al.
Pubblicazione: (2024)
ReCatcher: Towards LLMs Regression Testing for Code Generation
di: Abbassi, Altaf Allah, et al.
Pubblicazione: (2025)
di: Abbassi, Altaf Allah, et al.
Pubblicazione: (2025)
TaskEval: Assessing Difficulty of Code Generation Tasks for Large Language Models
di: Tambon, Florian, et al.
Pubblicazione: (2024)
di: Tambon, Florian, et al.
Pubblicazione: (2024)
Fault Localization in Deep Learning-based Software: A System-level Approach
di: Morovati, Mohammad Mehdi, et al.
Pubblicazione: (2024)
di: Morovati, Mohammad Mehdi, et al.
Pubblicazione: (2024)
Toward Debugging Deep Reinforcement Learning Programs with RLExplorer
di: Bouchoucha, Rached, et al.
Pubblicazione: (2024)
di: Bouchoucha, Rached, et al.
Pubblicazione: (2024)
An Efficient Model Maintenance Approach for MLOps
di: Majidi, Forough, et al.
Pubblicazione: (2024)
di: Majidi, Forough, et al.
Pubblicazione: (2024)
Bugs in Large Language Models Generated Code: An Empirical Study
di: Tambon, Florian, et al.
Pubblicazione: (2024)
di: Tambon, Florian, et al.
Pubblicazione: (2024)
Machine Learning Robustness: A Primer
di: Braiek, Houssem Ben, et al.
Pubblicazione: (2024)
di: Braiek, Houssem Ben, et al.
Pubblicazione: (2024)
What Information Contributes to Log-based Anomaly Detection? Insights from a Configurable Transformer-Based Approach
di: Wu, Xingfang, et al.
Pubblicazione: (2024)
di: Wu, Xingfang, et al.
Pubblicazione: (2024)
Imitation Game: Reproducing Deep Learning Bugs Leveraging an Intelligent Agent
di: Shah, Mehil B, et al.
Pubblicazione: (2025)
di: Shah, Mehil B, et al.
Pubblicazione: (2025)
Chain of Targeted Verification Questions to Improve the Reliability of Code Generated by LLMs
di: Ngassom, Sylvain Kouemo, et al.
Pubblicazione: (2024)
di: Ngassom, Sylvain Kouemo, et al.
Pubblicazione: (2024)
A Taxonomy of Inefficiencies in LLM-Generated Python Code
di: Abbassi, Altaf Allah, et al.
Pubblicazione: (2025)
di: Abbassi, Altaf Allah, et al.
Pubblicazione: (2025)
Quality Issues in Machine Learning Software Systems
di: Côté, Pierre-Olivier, et al.
Pubblicazione: (2023)
di: Côté, Pierre-Olivier, et al.
Pubblicazione: (2023)
LLMs and Stack Overflow Discussions: Reliability, Impact, and Challenges
di: Da Silva, Leuson, et al.
Pubblicazione: (2024)
di: Da Silva, Leuson, et al.
Pubblicazione: (2024)
A Survey of Bugs in AI-Generated Code
di: Gao, Ruofan, et al.
Pubblicazione: (2025)
di: Gao, Ruofan, et al.
Pubblicazione: (2025)
Drawing Pandas: A Benchmark for LLMs in Generating Plotting Code
di: Galimzyanov, Timur, et al.
Pubblicazione: (2024)
di: Galimzyanov, Timur, et al.
Pubblicazione: (2024)
PathOCL: Path-Based Prompt Augmentation for OCL Generation with GPT-4
di: Abukhalaf, Seif, et al.
Pubblicazione: (2024)
di: Abukhalaf, Seif, et al.
Pubblicazione: (2024)
GIST: Generated Inputs Sets Transferability in Deep Learning
di: Tambon, Florian, et al.
Pubblicazione: (2023)
di: Tambon, Florian, et al.
Pubblicazione: (2023)
RefAgent: A Multi-agent LLM-based Framework for Automatic Software Refactoring
di: Oueslati, Khouloud, et al.
Pubblicazione: (2025)
di: Oueslati, Khouloud, et al.
Pubblicazione: (2025)
Automating Code Adaptation for MLOps -- A Benchmarking Study on LLMs
di: Patel, Harsh, et al.
Pubblicazione: (2024)
di: Patel, Harsh, et al.
Pubblicazione: (2024)
Operational Robustness of LLMs on Code Generation
di: Paul, Debalina Ghosh, et al.
Pubblicazione: (2026)
di: Paul, Debalina Ghosh, et al.
Pubblicazione: (2026)
CodeTaste: Can LLMs Generate Human-Level Code Refactorings?
di: Thillen, Alex, et al.
Pubblicazione: (2026)
di: Thillen, Alex, et al.
Pubblicazione: (2026)
On the Effectiveness of Log Representation for Log-based Anomaly Detection
di: Wu, Xingfang, et al.
Pubblicazione: (2023)
di: Wu, Xingfang, et al.
Pubblicazione: (2023)
Protocode: Prototype-Driven Interpretability for Code Generation in LLMs
di: Bodla, Krishna Vamshi, et al.
Pubblicazione: (2025)
di: Bodla, Krishna Vamshi, et al.
Pubblicazione: (2025)
Deep-Bench: Deep Learning Benchmark Dataset for Code Generation
di: Daghighfarsoodeh, Alireza, et al.
Pubblicazione: (2025)
di: Daghighfarsoodeh, Alireza, et al.
Pubblicazione: (2025)
LiCoEval: Evaluating LLMs on License Compliance in Code Generation
di: Xu, Weiwei, et al.
Pubblicazione: (2024)
di: Xu, Weiwei, et al.
Pubblicazione: (2024)
Efficiency vs. Alignment: Investigating Safety and Fairness Risks in Parameter-Efficient Fine-Tuning of LLMs
di: Taraghi, Mina, et al.
Pubblicazione: (2025)
di: Taraghi, Mina, et al.
Pubblicazione: (2025)
VeriContest: A Competitive-Programming Benchmark for Verifiable Code Generation
di: Xie, Zichen, et al.
Pubblicazione: (2026)
di: Xie, Zichen, et al.
Pubblicazione: (2026)
SDLog: A Deep Learning Framework for Detecting Sensitive Information in Software Logs
di: Aghili, Roozbeh, et al.
Pubblicazione: (2025)
di: Aghili, Roozbeh, et al.
Pubblicazione: (2025)
CodeGeeX: A Pre-Trained Model for Code Generation with Multilingual Benchmarking on HumanEval-X
di: Zheng, Qinkai, et al.
Pubblicazione: (2023)
di: Zheng, Qinkai, et al.
Pubblicazione: (2023)
DevBench: A Realistic, Developer-Informed Benchmark for Code Generation Models
di: Kumarappan, Adarsh, et al.
Pubblicazione: (2026)
di: Kumarappan, Adarsh, et al.
Pubblicazione: (2026)
How Efficient is LLM-Generated Code? A Rigorous & High-Standard Benchmark
di: Qiu, Ruizhong, et al.
Pubblicazione: (2024)
di: Qiu, Ruizhong, et al.
Pubblicazione: (2024)
On LLMs' Internal Representation of Code Correctness
di: Ribeiro, Francisco, et al.
Pubblicazione: (2025)
di: Ribeiro, Francisco, et al.
Pubblicazione: (2025)
Common Challenges of Deep Reinforcement Learning Applications Development: An Empirical Study
di: Morovati, Mohammad Mehdi, et al.
Pubblicazione: (2023)
di: Morovati, Mohammad Mehdi, et al.
Pubblicazione: (2023)
Continuously Learning Bug Locations
di: Mindom, Paulina Stevia Nouwou, et al.
Pubblicazione: (2024)
di: Mindom, Paulina Stevia Nouwou, et al.
Pubblicazione: (2024)
miniCodeProps: a Minimal Benchmark for Proving Code Properties
di: Lohn, Evan, et al.
Pubblicazione: (2024)
di: Lohn, Evan, et al.
Pubblicazione: (2024)
Beyond Synthetic Benchmarks: Evaluating LLM Performance on Real-World Class-Level Code Generation
di: Rahman, Musfiqur, et al.
Pubblicazione: (2025)
di: Rahman, Musfiqur, et al.
Pubblicazione: (2025)
Evaluating the Use of LLMs for Documentation to Code Traceability
di: Alor, Ebube, et al.
Pubblicazione: (2025)
di: Alor, Ebube, et al.
Pubblicazione: (2025)
FairFLRep: Fairness aware fault localization and repair of Deep Neural Networks
di: Openja, Moses, et al.
Pubblicazione: (2025)
di: Openja, Moses, et al.
Pubblicazione: (2025)
Documenti analoghi
-
DeepCodeProbe: Towards Understanding What Models Trained on Code Learn
di: Majdinasab, Vahid, et al.
Pubblicazione: (2024) -
Trained Without My Consent: Detecting Code Inclusion In Language Models Trained on Code
di: Majdinasab, Vahid, et al.
Pubblicazione: (2024) -
ReCatcher: Towards LLMs Regression Testing for Code Generation
di: Abbassi, Altaf Allah, et al.
Pubblicazione: (2025) -
TaskEval: Assessing Difficulty of Code Generation Tasks for Large Language Models
di: Tambon, Florian, et al.
Pubblicazione: (2024) -
Fault Localization in Deep Learning-based Software: A System-level Approach
di: Morovati, Mohammad Mehdi, et al.
Pubblicazione: (2024)