Accuracy, Stability, and Repeated-Run Reliability of Large Language Models on Deterministic Programming Tasks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhou, Yongxi, Choi, Lai Yun, Wen, Jiaxi, Ye, Wenbo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Evaluating Reliability Gaps in Large Language Model Safety via Repeated Prompt Sampling
von: Broadwater, Keita
Veröffentlicht: (2026)
von: Broadwater, Keita
Veröffentlicht: (2026)
Automatic Programming: Large Language Models and Beyond
von: Lyu, Michael R., et al.
Veröffentlicht: (2024)
von: Lyu, Michael R., et al.
Veröffentlicht: (2024)
Benchmarking Large Language Models with Integer Sequence Generation Tasks
von: O'Malley, Daniel, et al.
Veröffentlicht: (2024)
von: O'Malley, Daniel, et al.
Veröffentlicht: (2024)
Toward Explaining Large Language Models in Software Engineering Tasks
von: Vitale, Antonio, et al.
Veröffentlicht: (2025)
von: Vitale, Antonio, et al.
Veröffentlicht: (2025)
Beyond Output Correctness: Benchmarking and Evaluating Large Language Model Reasoning in Coding Tasks
von: Li, Yuangang, et al.
Veröffentlicht: (2026)
von: Li, Yuangang, et al.
Veröffentlicht: (2026)
Code Red! On the Harmfulness of Applying Off-the-shelf Large Language Models to Programming Tasks
von: Al-Kaswan, Ali, et al.
Veröffentlicht: (2025)
von: Al-Kaswan, Ali, et al.
Veröffentlicht: (2025)
EsoLang-Bench: Evaluating Genuine Reasoning in Large Language Models via Esoteric Programming Languages
von: Sharma, Aman, et al.
Veröffentlicht: (2026)
von: Sharma, Aman, et al.
Veröffentlicht: (2026)
Assessing Large Language Models for Automated Feedback Generation in Learning Programming Problem Solving
von: Silva, Priscylla, et al.
Veröffentlicht: (2025)
von: Silva, Priscylla, et al.
Veröffentlicht: (2025)
Task Abstention for Large Language Models in Code Generation
von: Zhou, Yanke, et al.
Veröffentlicht: (2026)
von: Zhou, Yanke, et al.
Veröffentlicht: (2026)
Beyond Accuracy: Characterizing Code Comprehension Capabilities in (Large) Language Models
von: Mächtle, Felix, et al.
Veröffentlicht: (2026)
von: Mächtle, Felix, et al.
Veröffentlicht: (2026)
EduBot -- Can LLMs Solve Personalized Learning and Programming Assignments?
von: Wang, Yibin, et al.
Veröffentlicht: (2025)
von: Wang, Yibin, et al.
Veröffentlicht: (2025)
ModiGen: A Large Language Model-Based Workflow for Multi-Task Modelica Code Generation
von: Xiang, Jiahui, et al.
Veröffentlicht: (2025)
von: Xiang, Jiahui, et al.
Veröffentlicht: (2025)
Planning-Driven Programming: A Large Language Model Programming Workflow
von: Lei, Chao, et al.
Veröffentlicht: (2024)
von: Lei, Chao, et al.
Veröffentlicht: (2024)
TaskEval: Assessing Difficulty of Code Generation Tasks for Large Language Models
von: Tambon, Florian, et al.
Veröffentlicht: (2024)
von: Tambon, Florian, et al.
Veröffentlicht: (2024)
A Hierarchical Imprecise Probability Approach to Reliability Assessment of Large Language Models
von: Aghazadeh-Chakherlou, Robab, et al.
Veröffentlicht: (2025)
von: Aghazadeh-Chakherlou, Robab, et al.
Veröffentlicht: (2025)
StackSight: Unveiling WebAssembly through Large Language Models and Neurosymbolic Chain-of-Thought Decompilation
von: Fang, Weike, et al.
Veröffentlicht: (2024)
von: Fang, Weike, et al.
Veröffentlicht: (2024)
CodeFuse-13B: A Pretrained Multi-lingual Code Large Language Model
von: Di, Peng, et al.
Veröffentlicht: (2023)
von: Di, Peng, et al.
Veröffentlicht: (2023)
Generating Streamlining Constraints with Large Language Models
von: Voboril, Florentina, et al.
Veröffentlicht: (2024)
von: Voboril, Florentina, et al.
Veröffentlicht: (2024)
Are Large Language Models Memorizing Bug Benchmarks?
von: Ramos, Daniel, et al.
Veröffentlicht: (2024)
von: Ramos, Daniel, et al.
Veröffentlicht: (2024)
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs
von: Zhou, Zenghui, et al.
Veröffentlicht: (2026)
von: Zhou, Zenghui, et al.
Veröffentlicht: (2026)
Unlock the Correlation between Supervised Fine-Tuning and Reinforcement Learning in Training Code Large Language Models
von: Chen, Jie, et al.
Veröffentlicht: (2024)
von: Chen, Jie, et al.
Veröffentlicht: (2024)
The Impact of Fine-tuning Large Language Models on Automated Program Repair
von: Macháček, Roman, et al.
Veröffentlicht: (2025)
von: Macháček, Roman, et al.
Veröffentlicht: (2025)
On Integrating Large Language Models and Scenario-Based Programming for Improving Software Reliability
von: Berzack, Ayelet, et al.
Veröffentlicht: (2025)
von: Berzack, Ayelet, et al.
Veröffentlicht: (2025)
Beyond Accuracy: Policy Invariance as a Reliability Test for LLM Safety Judges
von: Weng, Shihao, et al.
Veröffentlicht: (2026)
von: Weng, Shihao, et al.
Veröffentlicht: (2026)
CrossPL: Evaluating Large Language Models on Cross Programming Language Code Generation
von: Xiong, Zhanhang, et al.
Veröffentlicht: (2025)
von: Xiong, Zhanhang, et al.
Veröffentlicht: (2025)
A Contemporary Survey of Large Language Model Assisted Program Analysis
von: Wang, Jiayimei, et al.
Veröffentlicht: (2025)
von: Wang, Jiayimei, et al.
Veröffentlicht: (2025)
Strategic Optimization and Challenges of Large Language Models in Object-Oriented Programming
von: Wang, Zinan
Veröffentlicht: (2024)
von: Wang, Zinan
Veröffentlicht: (2024)
Assessing the Impact of Code Changes on the Fault Localizability of Large Language Models
von: Haroon, Sabaat, et al.
Veröffentlicht: (2025)
von: Haroon, Sabaat, et al.
Veröffentlicht: (2025)
Exploring the Integration of Large Language Models in Industrial Test Maintenance Processes
von: Liu, Jingxiong, et al.
Veröffentlicht: (2024)
von: Liu, Jingxiong, et al.
Veröffentlicht: (2024)
ProToken: Token-Level Attribution for Federated Large Language Models
von: Gill, Waris, et al.
Veröffentlicht: (2026)
von: Gill, Waris, et al.
Veröffentlicht: (2026)
Effective Large Language Model Debugging with Best-first Tree Search
von: Song, Jialin, et al.
Veröffentlicht: (2024)
von: Song, Jialin, et al.
Veröffentlicht: (2024)
A-ProS: Towards Reliable Autonomous Programming Through Multi-Model Feedback
von: Tabassum, Anika, et al.
Veröffentlicht: (2026)
von: Tabassum, Anika, et al.
Veröffentlicht: (2026)
Evaluating Robustness of Large Language Models in Enterprise Applications: Benchmarks for Perturbation Consistency Across Formats and Languages
von: Bogavelli, Tara, et al.
Veröffentlicht: (2026)
von: Bogavelli, Tara, et al.
Veröffentlicht: (2026)
LLM-as-a-Judge for Scalable Test Coverage Evaluation: Accuracy, Operational Reliability, and Cost
von: Huang, Donghao, et al.
Veröffentlicht: (2025)
von: Huang, Donghao, et al.
Veröffentlicht: (2025)
RepoDebug: Repository-Level Multi-Task and Multi-Language Debugging Evaluation of Large Language Models
von: Liu, Jingjing, et al.
Veröffentlicht: (2025)
von: Liu, Jingjing, et al.
Veröffentlicht: (2025)
Data Wrangling Task Automation Using Code-Generating Language Models
von: Akella, Ashlesha, et al.
Veröffentlicht: (2025)
von: Akella, Ashlesha, et al.
Veröffentlicht: (2025)
Consolidating TinyML Lifecycle with Large Language Models: Reality, Illusion, or Opportunity?
von: Wu, Guanghan, et al.
Veröffentlicht: (2025)
von: Wu, Guanghan, et al.
Veröffentlicht: (2025)
Zero-Shot Attribution for Large Language Models: A Distribution Testing Approach
von: Canonne, Clément L., et al.
Veröffentlicht: (2025)
von: Canonne, Clément L., et al.
Veröffentlicht: (2025)
AKD : Adversarial Knowledge Distillation For Large Language Models Alignment on Coding tasks
von: Oulkadda, Ilyas, et al.
Veröffentlicht: (2025)
von: Oulkadda, Ilyas, et al.
Veröffentlicht: (2025)
CONSTRUCTA: Automating Commercial Construction Schedules in Fabrication Facilities with Large Language Models
von: Zhang, Yifan, et al.
Veröffentlicht: (2025)
von: Zhang, Yifan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Evaluating Reliability Gaps in Large Language Model Safety via Repeated Prompt Sampling
von: Broadwater, Keita
Veröffentlicht: (2026) -
Automatic Programming: Large Language Models and Beyond
von: Lyu, Michael R., et al.
Veröffentlicht: (2024) -
Benchmarking Large Language Models with Integer Sequence Generation Tasks
von: O'Malley, Daniel, et al.
Veröffentlicht: (2024) -
Toward Explaining Large Language Models in Software Engineering Tasks
von: Vitale, Antonio, et al.
Veröffentlicht: (2025) -
Beyond Output Correctness: Benchmarking and Evaluating Large Language Model Reasoning in Coding Tasks
von: Li, Yuangang, et al.
Veröffentlicht: (2026)