An Empirical Evaluation of Locally Deployed LLMs for Bug Detection in Python Code
Fuente:
arXiv
Saved in:
| Main Author: | Vulićević, Jelena Ilić |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
scicode-lint: Detecting Methodology Bugs in Scientific Python Code with LLM-Generated Patterns
by: Samsonau, Sergey V.
Published: (2026)
by: Samsonau, Sergey V.
Published: (2026)
On the Impact of Code Comments for Automated Bug-Fixing: An Empirical Study
by: Vitale, Antonio, et al.
Published: (2026)
by: Vitale, Antonio, et al.
Published: (2026)
Are Sparse Autoencoders Useful for Java Function Bug Detection?
by: Melo, Rui, et al.
Published: (2025)
by: Melo, Rui, et al.
Published: (2025)
Evaluating the Use of LLMs for Documentation to Code Traceability
by: Alor, Ebube, et al.
Published: (2025)
by: Alor, Ebube, et al.
Published: (2025)
SWT-Bench: Testing and Validating Real-World Bug-Fixes with Code Agents
by: Mündler, Niels, et al.
Published: (2024)
by: Mündler, Niels, et al.
Published: (2024)
The Struggles of LLMs in Cross-lingual Code Clone Detection
by: Moumoula, Micheline Bénédicte, et al.
Published: (2024)
by: Moumoula, Micheline Bénédicte, et al.
Published: (2024)
GREPO: A Benchmark for Graph Neural Networks on Repository-Level Bug Localization
by: Wang, Juntong, et al.
Published: (2026)
by: Wang, Juntong, et al.
Published: (2026)
Can LLMs Find Bugs in Code? An Evaluation from Beginner Errors to Security Vulnerabilities in Python and C++
by: Mhatre, Akshay, et al.
Published: (2025)
by: Mhatre, Akshay, et al.
Published: (2025)
SnipGen: A Mining Repository Framework for Evaluating LLMs for Code
by: Rodriguez-Cardenas, Daniel, et al.
Published: (2025)
by: Rodriguez-Cardenas, Daniel, et al.
Published: (2025)
LiCoEval: Evaluating LLMs on License Compliance in Code Generation
by: Xu, Weiwei, et al.
Published: (2024)
by: Xu, Weiwei, et al.
Published: (2024)
Are Large Language Models Memorizing Bug Benchmarks?
by: Ramos, Daniel, et al.
Published: (2024)
by: Ramos, Daniel, et al.
Published: (2024)
Reducing False Positives in Static Bug Detection with LLMs: An Empirical Study in Industry
by: Du, Xueying, et al.
Published: (2026)
by: Du, Xueying, et al.
Published: (2026)
Operational Robustness of LLMs on Code Generation
by: Paul, Debalina Ghosh, et al.
Published: (2026)
by: Paul, Debalina Ghosh, et al.
Published: (2026)
On LLMs' Internal Representation of Code Correctness
by: Ribeiro, Francisco, et al.
Published: (2025)
by: Ribeiro, Francisco, et al.
Published: (2025)
CodeTaste: Can LLMs Generate Human-Level Code Refactorings?
by: Thillen, Alex, et al.
Published: (2026)
by: Thillen, Alex, et al.
Published: (2026)
Can LLMs Generate Architectural Design Decisions? -An Exploratory Empirical study
by: Dhar, Rudra, et al.
Published: (2024)
by: Dhar, Rudra, et al.
Published: (2024)
Towards Evaluation Engineering: An Empirical Study of ML Evaluation Harnesses in the Wild
by: Zhao, Zhimin, et al.
Published: (2026)
by: Zhao, Zhimin, et al.
Published: (2026)
More with Less: An Empirical Study of Turn-Control Strategies for Efficient Coding Agents
by: Gao, Pengfei, et al.
Published: (2025)
by: Gao, Pengfei, et al.
Published: (2025)
LLMs in Coding and their Impact on the Commercial Software Engineering Landscape
by: Belozerov, Vladislav, et al.
Published: (2025)
by: Belozerov, Vladislav, et al.
Published: (2025)
Protocode: Prototype-Driven Interpretability for Code Generation in LLMs
by: Bodla, Krishna Vamshi, et al.
Published: (2025)
by: Bodla, Krishna Vamshi, et al.
Published: (2025)
Imitation Game: Reproducing Deep Learning Bugs Leveraging an Intelligent Agent
by: Shah, Mehil B, et al.
Published: (2025)
by: Shah, Mehil B, et al.
Published: (2025)
Coding in a Bubble? Evaluating LLMs in Resolving Context Adaptation Bugs During Code Adaptation
by: Zhang, Tanghaoran, et al.
Published: (2026)
by: Zhang, Tanghaoran, et al.
Published: (2026)
Bugs in Large Language Models Generated Code: An Empirical Study
by: Tambon, Florian, et al.
Published: (2024)
by: Tambon, Florian, et al.
Published: (2024)
Towards a Neural Debugger for Python
by: Beck, Maximilian, et al.
Published: (2026)
by: Beck, Maximilian, et al.
Published: (2026)
Automating Code Adaptation for MLOps -- A Benchmarking Study on LLMs
by: Patel, Harsh, et al.
Published: (2024)
by: Patel, Harsh, et al.
Published: (2024)
Drawing Pandas: A Benchmark for LLMs in Generating Plotting Code
by: Galimzyanov, Timur, et al.
Published: (2024)
by: Galimzyanov, Timur, et al.
Published: (2024)
Free and Customizable Code Documentation with LLMs: A Fine-Tuning Approach
by: Chakrabarty, Sayak, et al.
Published: (2024)
by: Chakrabarty, Sayak, et al.
Published: (2024)
TriagerX: Dual Transformers for Bug Triaging Tasks with Content and Interaction Based Rankings
by: Mamun, Md Afif Al, et al.
Published: (2025)
by: Mamun, Md Afif Al, et al.
Published: (2025)
Order Matters! An Empirical Study on Large Language Models' Input Order Bias in Software Fault Localization
by: Rafi, Md Nakhla, et al.
Published: (2024)
by: Rafi, Md Nakhla, et al.
Published: (2024)
Evaluation of LLMs on Syntax-Aware Code Fill-in-the-Middle Tasks
by: Gong, Linyuan, et al.
Published: (2024)
by: Gong, Linyuan, et al.
Published: (2024)
Towards More Trustworthy and Interpretable LLMs for Code through Syntax-Grounded Explanations
by: Palacio, David N., et al.
Published: (2024)
by: Palacio, David N., et al.
Published: (2024)
Prism: Dynamic and Flexible Benchmarking of LLMs Code Generation with Monte Carlo Tree Search
by: Majdinasab, Vahid, et al.
Published: (2025)
by: Majdinasab, Vahid, et al.
Published: (2025)
Keeping Code-Aware LLMs Fresh: Full Refresh, In-Context Deltas, and Incremental Fine-Tuning
by: Sharma, Pradeep Kumar, et al.
Published: (2025)
by: Sharma, Pradeep Kumar, et al.
Published: (2025)
Can We Enhance Bug Report Quality Using LLMs?: An Empirical Study of LLM-Based Bug Report Generation
by: Acharya, Jagrit, et al.
Published: (2025)
by: Acharya, Jagrit, et al.
Published: (2025)
Benchmarking Reward Hack Detection in Code Environments via Contrastive Analysis
by: Deshpande, Darshan, et al.
Published: (2026)
by: Deshpande, Darshan, et al.
Published: (2026)
SCoPE: Evaluating LLMs for Software Vulnerability Detection
by: Gonçalves, José, et al.
Published: (2024)
by: Gonçalves, José, et al.
Published: (2024)
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs
by: Zhou, Zenghui, et al.
Published: (2026)
by: Zhou, Zenghui, et al.
Published: (2026)
Deploying Privacy Guardrails for LLMs: A Comparative Analysis of Real-World Applications
by: Asthana, Shubhi, et al.
Published: (2025)
by: Asthana, Shubhi, et al.
Published: (2025)
Model See, Model Do? Exposure-Aware Evaluation of Bug-vs-Fix Preference in Code LLMs
by: Al-Kaswan, Ali, et al.
Published: (2026)
by: Al-Kaswan, Ali, et al.
Published: (2026)
AIPC: Agent-Based Automation for AI Model Deployment with Qualcomm AI Runtime
by: Su, Jianhao, et al.
Published: (2026)
by: Su, Jianhao, et al.
Published: (2026)
Similar Items
-
scicode-lint: Detecting Methodology Bugs in Scientific Python Code with LLM-Generated Patterns
by: Samsonau, Sergey V.
Published: (2026) -
On the Impact of Code Comments for Automated Bug-Fixing: An Empirical Study
by: Vitale, Antonio, et al.
Published: (2026) -
Are Sparse Autoencoders Useful for Java Function Bug Detection?
by: Melo, Rui, et al.
Published: (2025) -
Evaluating the Use of LLMs for Documentation to Code Traceability
by: Alor, Ebube, et al.
Published: (2025) -
SWT-Bench: Testing and Validating Real-World Bug-Fixes with Code Agents
by: Mündler, Niels, et al.
Published: (2024)