Adaptive Hierarchical Evaluation of LLMs and SAST tools for CWE Prediction in Python
Fuente:
arXiv
Salvato in:
| Autori principali: | Adnan, Muntasir, Kuhn, Carlos C. N. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
The Debugging Decay Index: Rethinking Debugging Strategies for Code LLMs
di: Adnan, Muntasir, et al.
Pubblicazione: (2025)
di: Adnan, Muntasir, et al.
Pubblicazione: (2025)
Large Language Model Guided Self-Debugging Code Generation
di: Adnan, Muntasir, et al.
Pubblicazione: (2025)
di: Adnan, Muntasir, et al.
Pubblicazione: (2025)
CASTLE: Benchmarking Dataset for Static Code Analyzers and LLMs towards CWE Detection
di: Dubniczky, Richard A., et al.
Pubblicazione: (2025)
di: Dubniczky, Richard A., et al.
Pubblicazione: (2025)
The Last Dependency Crusade: Solving Python Dependency Conflicts with LLMs
di: Bartlett, Antony, et al.
Pubblicazione: (2025)
di: Bartlett, Antony, et al.
Pubblicazione: (2025)
An Empirical Evaluation of Locally Deployed LLMs for Bug Detection in Python Code
di: Vulićević, Jelena Ilić
Pubblicazione: (2026)
di: Vulićević, Jelena Ilić
Pubblicazione: (2026)
HCAG: Hierarchical Abstraction and Retrieval-Augmented Generation on Theoretical Repositories with LLMs
di: Wu, Yusen, et al.
Pubblicazione: (2026)
di: Wu, Yusen, et al.
Pubblicazione: (2026)
Evaluating AI-generated code for C++, Fortran, Go, Java, Julia, Matlab, Python, R, and Rust
di: Diehl, Patrick, et al.
Pubblicazione: (2024)
di: Diehl, Patrick, et al.
Pubblicazione: (2024)
Evaluating LLMs for Visualization Tasks
di: Khan, Saadiq Rauf, et al.
Pubblicazione: (2025)
di: Khan, Saadiq Rauf, et al.
Pubblicazione: (2025)
Verification-Guided Context Optimization for Tool Calling via Hierarchical LLMs-as-Editors
di: Li, Henger, et al.
Pubblicazione: (2025)
di: Li, Henger, et al.
Pubblicazione: (2025)
Hierarchical Repository-Level Code Summarization for Business Applications Using Local LLMs
di: Dhulshette, Nilesh, et al.
Pubblicazione: (2025)
di: Dhulshette, Nilesh, et al.
Pubblicazione: (2025)
Better Python Programming for all: With the focus on Maintainability
di: Shivashankar, Karthik, et al.
Pubblicazione: (2024)
di: Shivashankar, Karthik, et al.
Pubblicazione: (2024)
Evaluating the Energy-Efficiency of the Code Generated by LLMs
di: Islam, Md Arman, et al.
Pubblicazione: (2025)
di: Islam, Md Arman, et al.
Pubblicazione: (2025)
Evaluating the Generalizability of LLMs in Automated Program Repair
di: Li, Fengjie, et al.
Pubblicazione: (2025)
di: Li, Fengjie, et al.
Pubblicazione: (2025)
VisCoder: Fine-Tuning LLMs for Executable Python Visualization Code Generation
di: Ni, Yuansheng, et al.
Pubblicazione: (2025)
di: Ni, Yuansheng, et al.
Pubblicazione: (2025)
Holistic Evaluation of State-of-the-Art LLMs for Code Generation
di: Zhang, Le, et al.
Pubblicazione: (2025)
di: Zhang, Le, et al.
Pubblicazione: (2025)
Using LLMs in Software Requirements Specifications: An Empirical Evaluation
di: Krishna, Madhava, et al.
Pubblicazione: (2024)
di: Krishna, Madhava, et al.
Pubblicazione: (2024)
Machine Learning Techniques for Python Source Code Vulnerability Detection
di: Farasat, Talaya, et al.
Pubblicazione: (2024)
di: Farasat, Talaya, et al.
Pubblicazione: (2024)
TAM-Eval: Evaluating LLMs for Automated Unit Test Maintenance
di: Bruches, Elena, et al.
Pubblicazione: (2026)
di: Bruches, Elena, et al.
Pubblicazione: (2026)
Benchmark Dataset Generation and Evaluation for Excel Formula Repair with LLMs
di: Singha, Ananya, et al.
Pubblicazione: (2025)
di: Singha, Ananya, et al.
Pubblicazione: (2025)
FullStack Bench: Evaluating LLMs as Full Stack Coders
di: Bytedance-Seed-Foundation-Code-Team, et al.
Pubblicazione: (2024)
di: Bytedance-Seed-Foundation-Code-Team, et al.
Pubblicazione: (2024)
Quality and Security Signals in AI-Generated Python Refactoring Pull Requests
di: Almukhtar, Mohamed, et al.
Pubblicazione: (2026)
di: Almukhtar, Mohamed, et al.
Pubblicazione: (2026)
FrontendBench: A Benchmark for Evaluating LLMs on Front-End Development via Automatic Evaluation
di: Zhu, Hongda, et al.
Pubblicazione: (2025)
di: Zhu, Hongda, et al.
Pubblicazione: (2025)
Agentic Property-Based Testing: Finding Bugs Across the Python Ecosystem
di: Maaz, Muhammad, et al.
Pubblicazione: (2025)
di: Maaz, Muhammad, et al.
Pubblicazione: (2025)
GBQA: A Game Benchmark for Evaluating LLMs as Quality Assurance Engineers
di: Jiang, Shufan, et al.
Pubblicazione: (2026)
di: Jiang, Shufan, et al.
Pubblicazione: (2026)
Evaluating the Effectiveness of LLMs in Fixing Maintainability Issues in Real-World Projects
di: Nunes, Henrique, et al.
Pubblicazione: (2025)
di: Nunes, Henrique, et al.
Pubblicazione: (2025)
Open the Oyster: Empirical Evaluation and Improvement of Code Reasoning Confidence in LLMs
di: Wang, Shufan, et al.
Pubblicazione: (2025)
di: Wang, Shufan, et al.
Pubblicazione: (2025)
Mutation-based Consistency Testing for Evaluating the Code Understanding Capability of LLMs
di: Li, Ziyu, et al.
Pubblicazione: (2024)
di: Li, Ziyu, et al.
Pubblicazione: (2024)
TREAT: A Code LLMs Trustworthiness / Reliability Evaluation and Testing Framework
di: Gao, Shuzheng, et al.
Pubblicazione: (2025)
di: Gao, Shuzheng, et al.
Pubblicazione: (2025)
Benchmarking Text-to-Python against Text-to-SQL: The Impact of Explicit Logic and Ambiguity
di: Hu, Hangle, et al.
Pubblicazione: (2026)
di: Hu, Hangle, et al.
Pubblicazione: (2026)
PyGen: A Collaborative Human-AI Approach to Python Package Creation
di: Barua, Saikat, et al.
Pubblicazione: (2024)
di: Barua, Saikat, et al.
Pubblicazione: (2024)
TypyBench: Evaluating LLM Type Inference for Untyped Python Repositories
di: Dong, Honghua, et al.
Pubblicazione: (2025)
di: Dong, Honghua, et al.
Pubblicazione: (2025)
Evaluating Human Trajectory Prediction with Metamorphic Testing
di: Spieker, Helge, et al.
Pubblicazione: (2024)
di: Spieker, Helge, et al.
Pubblicazione: (2024)
Evaluating the Use of LLMs for Automated DOM-Level Resolution of Web Performance Issues
di: Peters, Gideon, et al.
Pubblicazione: (2026)
di: Peters, Gideon, et al.
Pubblicazione: (2026)
PlotChain: Deterministic Checkpointed Evaluation of Multimodal LLMs on Engineering Plot Reading
di: Ravishankara, Mayank
Pubblicazione: (2026)
di: Ravishankara, Mayank
Pubblicazione: (2026)
MEMRES: A Memory-Augmented Resolver with Confidence Cascade for Agentic Python Dependency Resolution
di: Minh, Dao Sy Duy, et al.
Pubblicazione: (2026)
di: Minh, Dao Sy Duy, et al.
Pubblicazione: (2026)
From Empirical Evaluation to Context-Aware Enhancement: Repairing Regression Errors with LLMs
di: Ho, Anh, et al.
Pubblicazione: (2025)
di: Ho, Anh, et al.
Pubblicazione: (2025)
WebDevJudge: Evaluating (M)LLMs as Critiques for Web Development Quality
di: Li, Chunyang, et al.
Pubblicazione: (2025)
di: Li, Chunyang, et al.
Pubblicazione: (2025)
Copilot-in-the-Loop: Fixing Code Smells in Copilot-Generated Python Code using Copilot
di: Zhang, Beiqi, et al.
Pubblicazione: (2024)
di: Zhang, Beiqi, et al.
Pubblicazione: (2024)
Capture the Flags: Family-Based Evaluation of Agentic LLMs via Semantics-Preserving Transformations
di: Honarvar, Shahin, et al.
Pubblicazione: (2026)
di: Honarvar, Shahin, et al.
Pubblicazione: (2026)
PyVeritas: On Verifying Python via LLM-Based Transpilation and Bounded Model Checking for C
di: Orvalho, Pedro, et al.
Pubblicazione: (2025)
di: Orvalho, Pedro, et al.
Pubblicazione: (2025)
Documenti analoghi
-
The Debugging Decay Index: Rethinking Debugging Strategies for Code LLMs
di: Adnan, Muntasir, et al.
Pubblicazione: (2025) -
Large Language Model Guided Self-Debugging Code Generation
di: Adnan, Muntasir, et al.
Pubblicazione: (2025) -
CASTLE: Benchmarking Dataset for Static Code Analyzers and LLMs towards CWE Detection
di: Dubniczky, Richard A., et al.
Pubblicazione: (2025) -
The Last Dependency Crusade: Solving Python Dependency Conflicts with LLMs
di: Bartlett, Antony, et al.
Pubblicazione: (2025) -
An Empirical Evaluation of Locally Deployed LLMs for Bug Detection in Python Code
di: Vulićević, Jelena Ilić
Pubblicazione: (2026)