Evaluating the Limitations of Local LLMs in Solving Complex Programming Challenges
Fuente:
arXiv
Saved in:
| Main Authors: | Matotek, Kadin, Cassel, Heather, Amiruzzaman, Md, Ngo, Linh B. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Developer Challenges on Large Language Models: A Study of Stack Overflow and OpenAI Developer Forum Posts
by: Alam, Khairul, et al.
Published: (2024)
by: Alam, Khairul, et al.
Published: (2024)
Toward Architecture-Aware Evaluation Metrics for LLM Agents
by: Souza, Débora, et al.
Published: (2026)
by: Souza, Débora, et al.
Published: (2026)
AdaDec: A Uncertainty-Guided Lookahead Decoding Framework for LLM-Based Code Generation
by: He, Kaifeng, et al.
Published: (2025)
by: He, Kaifeng, et al.
Published: (2025)
Evaluating the efficacy of LLM Safety Solutions : The Palit Benchmark Dataset
by: Palit, Sayon, et al.
Published: (2025)
by: Palit, Sayon, et al.
Published: (2025)
RelRepair: Enhancing Automated Program Repair by Retrieving Relevant Code
by: Liu, Shunyu, et al.
Published: (2025)
by: Liu, Shunyu, et al.
Published: (2025)
Circularity and Symmetries of $p$ and $p^{2}$-polygons
by: Haag, Rolf
Published: (2025)
by: Haag, Rolf
Published: (2025)
Can AI Assist in Olympiad Coding
by: Ren, Samuel
Published: (2025)
by: Ren, Samuel
Published: (2025)
Mobile Phone Sensor-based Nigerian Driving Dataset to Detect Alcohol-influenced Behaviours
by: Thompson, Iniakpokeikiye Peter, et al.
Published: (2025)
by: Thompson, Iniakpokeikiye Peter, et al.
Published: (2025)
Random Heterogeneous Neurochaos Learning Architecture for Data Classification
by: S, Remya Ajai A, et al.
Published: (2024)
by: S, Remya Ajai A, et al.
Published: (2024)
PromptSAM+: Malware Detection based on Prompt Segment Anything Model
by: Wei, Xingyuan, et al.
Published: (2024)
by: Wei, Xingyuan, et al.
Published: (2024)
The Hidden Attention of Mamba Models
by: Ali, Ameen, et al.
Published: (2024)
by: Ali, Ameen, et al.
Published: (2024)
From Pixels to Privacy: Temporally Consistent Video Anonymization via Token Pruning for Privacy Preserving Action Recognition
by: Aslam, Nazia, et al.
Published: (2026)
by: Aslam, Nazia, et al.
Published: (2026)
Improved IR-based Bug Localization with Intelligent Relevance Feedback
by: Samir, Asif Mohammed, et al.
Published: (2025)
by: Samir, Asif Mohammed, et al.
Published: (2025)
Contrastive Learning-Enhanced Large Language Models for Monolith-to-Microservice Decomposition
by: Sellami, Khaled, et al.
Published: (2025)
by: Sellami, Khaled, et al.
Published: (2025)
Software Implementation of Digital Filtering via Tustin's Bilinear Transform
by: Herron, Connor W.
Published: (2024)
by: Herron, Connor W.
Published: (2024)
R-Genie: Reasoning-Guided Generative Image Editing
by: Zhang, Dong, et al.
Published: (2025)
by: Zhang, Dong, et al.
Published: (2025)
Train to Defend: First Defense Against Cryptanalytic Neural Network Parameter Extraction Attacks
by: Kurian, Ashley, et al.
Published: (2025)
by: Kurian, Ashley, et al.
Published: (2025)
CIFE: Code Instruction-Following Evaluation
by: Gunnu, Sravani, et al.
Published: (2025)
by: Gunnu, Sravani, et al.
Published: (2025)
Leveraging Test Driven Development with Large Language Models for Reliable and Verifiable Spreadsheet Code Generation: A Research Framework
by: Thorne, Simon, et al.
Published: (2025)
by: Thorne, Simon, et al.
Published: (2025)
AgentModernize: Preserving Business Logic in Legacy Modernization with Multi-Agent LLMs and Behavioral Specification Graphs
by: Ahmed, Sheikh Nazib, et al.
Published: (2026)
by: Ahmed, Sheikh Nazib, et al.
Published: (2026)
EvoGraph: Hybrid Directed Graph Evolution toward Software 3.0
by: Costa, Igor, et al.
Published: (2025)
by: Costa, Igor, et al.
Published: (2025)
OODEval: Evaluating Large Language Models on Object-Oriented Design
by: Xiao, Bingxu, et al.
Published: (2026)
by: Xiao, Bingxu, et al.
Published: (2026)
Only Whats Necessary: Pareto Optimal Data Minimization for Privacy Preserving Video Anomaly Detection
by: Aslam, Nazia, et al.
Published: (2026)
by: Aslam, Nazia, et al.
Published: (2026)
Plan with Code: Comparing approaches for robust NL to DSL generation
by: Bassamzadeh, Nastaran, et al.
Published: (2024)
by: Bassamzadeh, Nastaran, et al.
Published: (2024)
A Comparative Study of DSL Code Generation: Fine-Tuning vs. Optimized Retrieval Augmentation
by: Bassamzadeh, Nastaran, et al.
Published: (2024)
by: Bassamzadeh, Nastaran, et al.
Published: (2024)
Addressing Data Leakage in HumanEval Using Combinatorial Test Design
by: Bradbury, Jeremy S., et al.
Published: (2024)
by: Bradbury, Jeremy S., et al.
Published: (2024)
Compiled AI: Deterministic Code Generation for LLM-Based Workflow Automation
by: Trooskens, Geert, et al.
Published: (2026)
by: Trooskens, Geert, et al.
Published: (2026)
CBR -- Boosting Adaptive Classification By Retrieval of Encrypted Network Traffic with Out-of-distribution
by: Lukach, Amir, et al.
Published: (2024)
by: Lukach, Amir, et al.
Published: (2024)
Leveraging Large Language Models for Use Case Model Generation from Software Requirements
by: Eisenreich, Tobias, et al.
Published: (2025)
by: Eisenreich, Tobias, et al.
Published: (2025)
Automating Domain-Driven Design: Experience with a Prompting Framework
by: Eisenreich, Tobias, et al.
Published: (2026)
by: Eisenreich, Tobias, et al.
Published: (2026)
Towards Observation Lakehouses: Living, Interactive Archives of Software Behavior
by: Kessel, Marcus
Published: (2025)
by: Kessel, Marcus
Published: (2025)
A Detailed Comparative Analysis of Blockchain Consensus Mechanisms
by: Andrews, Kaeli, et al.
Published: (2025)
by: Andrews, Kaeli, et al.
Published: (2025)
Uplink Transmit Power Optimization for Distributed Massive MIMO Systems with 1-Bit ADCs
by: Gouda, Bikshapathi, et al.
Published: (2024)
by: Gouda, Bikshapathi, et al.
Published: (2024)
Semi-Supervised Radiomics for Glioblastoma IDH Mutation: Limited Labels, Data Sensitivity, and SHAP Interpretation
by: Pouria, Amir Hossein, et al.
Published: (2025)
by: Pouria, Amir Hossein, et al.
Published: (2025)
Idea2Story: An Automated Pipeline for Transforming Research Concepts into Complete Scientific Narratives
by: Xu, Tengyue, et al.
Published: (2026)
by: Xu, Tengyue, et al.
Published: (2026)
Dreaming Falcon: Physics-Informed Model-Based Reinforcement Learning for Quadcopters
by: Vytla, Eashan, et al.
Published: (2025)
by: Vytla, Eashan, et al.
Published: (2025)
AI-assisted 3D Preservation and Reconstruction of Temple Arts
by: Shih, Naai-Jung
Published: (2025)
by: Shih, Naai-Jung
Published: (2025)
Do Large Language Models Speak All Languages Equally? A Comparative Study in Low-Resource Settings
by: Hasan, Md. Arid, et al.
Published: (2024)
by: Hasan, Md. Arid, et al.
Published: (2024)
Synergy of Large Language Model and Model Driven Engineering for Automated Development of Centralized Vehicular Systems
by: Petrovic, Nenad, et al.
Published: (2024)
by: Petrovic, Nenad, et al.
Published: (2024)
Towards Single-System Illusion in Software-Defined Vehicles -- Automated, AI-Powered Workflow
by: Lebioda, Krzysztof, et al.
Published: (2024)
by: Lebioda, Krzysztof, et al.
Published: (2024)
Similar Items
-
Developer Challenges on Large Language Models: A Study of Stack Overflow and OpenAI Developer Forum Posts
by: Alam, Khairul, et al.
Published: (2024) -
Toward Architecture-Aware Evaluation Metrics for LLM Agents
by: Souza, Débora, et al.
Published: (2026) -
AdaDec: A Uncertainty-Guided Lookahead Decoding Framework for LLM-Based Code Generation
by: He, Kaifeng, et al.
Published: (2025) -
Evaluating the efficacy of LLM Safety Solutions : The Palit Benchmark Dataset
by: Palit, Sayon, et al.
Published: (2025) -
RelRepair: Enhancing Automated Program Repair by Retrieving Relevant Code
by: Liu, Shunyu, et al.
Published: (2025)