Towards a Probabilistic Framework for Analyzing and Improving LLM-Enabled Software
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Baldonado, Juan Manuel, Bonomo-Braberman, Flavia, Braberman, Víctor Adrián |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Generative transformations and patterns in LLM-native approaches for software verification and falsification
von: Braberman, Víctor A., et al.
Veröffentlicht: (2024)
von: Braberman, Víctor A., et al.
Veröffentlicht: (2024)
Automated structural testing of LLM-based agents: methods, framework, and case studies
von: Kohl, Jens, et al.
Veröffentlicht: (2026)
von: Kohl, Jens, et al.
Veröffentlicht: (2026)
AgentEval: DAG-Structured Step-Level Evaluation for Agentic Workflows with Error Propagation Tracking
von: Guo, Dongxin, et al.
Veröffentlicht: (2026)
von: Guo, Dongxin, et al.
Veröffentlicht: (2026)
AgentModernize: Preserving Business Logic in Legacy Modernization with Multi-Agent LLMs and Behavioral Specification Graphs
von: Ahmed, Sheikh Nazib, et al.
Veröffentlicht: (2026)
von: Ahmed, Sheikh Nazib, et al.
Veröffentlicht: (2026)
Towards Observation Lakehouses: Living, Interactive Archives of Software Behavior
von: Kessel, Marcus
Veröffentlicht: (2025)
von: Kessel, Marcus
Veröffentlicht: (2025)
RelRepair: Enhancing Automated Program Repair by Retrieving Relevant Code
von: Liu, Shunyu, et al.
Veröffentlicht: (2025)
von: Liu, Shunyu, et al.
Veröffentlicht: (2025)
An Analysis of LLM Fine-Tuning and Few-Shot Learning for Flaky Test Detection and Classification
von: More, Riddhi, et al.
Veröffentlicht: (2025)
von: More, Riddhi, et al.
Veröffentlicht: (2025)
When Retrieval Hurts Code Completion: A Diagnostic Study of Stale Repository Context
von: Weng, Haojun, et al.
Veröffentlicht: (2026)
von: Weng, Haojun, et al.
Veröffentlicht: (2026)
Reducing Maintenance Burden in Behaviour-Driven Development: A Paraphrase-Robust Duplicate-Step Detector with a 1.1M-Step Open Benchmark
von: Mughal, Ali Hassaan, et al.
Veröffentlicht: (2026)
von: Mughal, Ali Hassaan, et al.
Veröffentlicht: (2026)
GALA: Multimodal Graph Alignment for Bug Localization in Automated Program Repair
von: Liu, Zhuoyao, et al.
Veröffentlicht: (2026)
von: Liu, Zhuoyao, et al.
Veröffentlicht: (2026)
Mining Subscenario Refactoring Opportunities in Behaviour-Driven Software Test Suites: ML Classifiers and LLM-Judge Baselines
von: Mughal, Ali Hassaan, et al.
Veröffentlicht: (2026)
von: Mughal, Ali Hassaan, et al.
Veröffentlicht: (2026)
From Untestable to Testable: Metamorphic Testing in the Age of LLMs
von: Terragni, Valerio
Veröffentlicht: (2026)
von: Terragni, Valerio
Veröffentlicht: (2026)
Assessing Data Augmentation-Induced Bias in Training and Testing of Machine Learning Models
von: More, Riddhi, et al.
Veröffentlicht: (2025)
von: More, Riddhi, et al.
Veröffentlicht: (2025)
A Systematic Approach for Assessing Large Language Models' Test Case Generation Capability
von: Chang, Hung-Fu, et al.
Veröffentlicht: (2025)
von: Chang, Hung-Fu, et al.
Veröffentlicht: (2025)
Test-driven Software Experimentation with LASSO: an LLM Prompt Benchmarking Example
von: Kessel, Marcus
Veröffentlicht: (2024)
von: Kessel, Marcus
Veröffentlicht: (2024)
The Kieker Observability Framework Version 2
von: Yang, Shinhyung, et al.
Veröffentlicht: (2025)
von: Yang, Shinhyung, et al.
Veröffentlicht: (2025)
FORGE: An LLM-driven Framework for Large-Scale Smart Contract Vulnerability Dataset Construction
von: Chen, Jiachi, et al.
Veröffentlicht: (2025)
von: Chen, Jiachi, et al.
Veröffentlicht: (2025)
CIFE: Code Instruction-Following Evaluation
von: Gunnu, Sravani, et al.
Veröffentlicht: (2025)
von: Gunnu, Sravani, et al.
Veröffentlicht: (2025)
Addressing Data Leakage in HumanEval Using Combinatorial Test Design
von: Bradbury, Jeremy S., et al.
Veröffentlicht: (2024)
von: Bradbury, Jeremy S., et al.
Veröffentlicht: (2024)
Secure coding for web applications: Frameworks, challenges, and the role of LLMs
von: Kiashemshaki, Kiana, et al.
Veröffentlicht: (2025)
von: Kiashemshaki, Kiana, et al.
Veröffentlicht: (2025)
Toward Architecture-Aware Evaluation Metrics for LLM Agents
von: Souza, Débora, et al.
Veröffentlicht: (2026)
von: Souza, Débora, et al.
Veröffentlicht: (2026)
Morescient GAI for Software Engineering (Extended Version)
von: Kessel, Marcus, et al.
Veröffentlicht: (2024)
von: Kessel, Marcus, et al.
Veröffentlicht: (2024)
TDAD: Test-Driven Agentic Development - Reducing Code Regressions in AI Coding Agents via Graph-Based Impact Analysis
von: Alonso, Pepe, et al.
Veröffentlicht: (2026)
von: Alonso, Pepe, et al.
Veröffentlicht: (2026)
AuditRepairBench: A Paired-Execution Trace Corpus for Evaluator-Channel Ranking Instability in Agent Repair
von: Hu, Yuelin, et al.
Veröffentlicht: (2026)
von: Hu, Yuelin, et al.
Veröffentlicht: (2026)
CoTran: An LLM-based Code Translator using Reinforcement Learning with Feedback from Compiler and Symbolic Execution
von: Jana, Prithwish, et al.
Veröffentlicht: (2023)
von: Jana, Prithwish, et al.
Veröffentlicht: (2023)
Technical Debt Management: The Road Ahead for Successful Software Delivery
von: Avgeriou, Paris, et al.
Veröffentlicht: (2024)
von: Avgeriou, Paris, et al.
Veröffentlicht: (2024)
SmellBench: Evaluating LLM Agents on Architectural Code Smell Repair
von: Dinu, Ion George, et al.
Veröffentlicht: (2026)
von: Dinu, Ion George, et al.
Veröffentlicht: (2026)
The Specification as Quality Gate: Three Hypotheses on AI-Assisted Code Review
von: Zietsman, Christo
Veröffentlicht: (2026)
von: Zietsman, Christo
Veröffentlicht: (2026)
Rango: Adaptive Retrieval-Augmented Proving for Automated Software Verification
von: Thompson, Kyle, et al.
Veröffentlicht: (2024)
von: Thompson, Kyle, et al.
Veröffentlicht: (2024)
Towards Single-System Illusion in Software-Defined Vehicles -- Automated, AI-Powered Workflow
von: Lebioda, Krzysztof, et al.
Veröffentlicht: (2024)
von: Lebioda, Krzysztof, et al.
Veröffentlicht: (2024)
LLM-Assisted Translation of Legacy FORTRAN Codes to C++: A Cross-Platform Study
von: Ranasinghe, Nishath Rajiv, et al.
Veröffentlicht: (2025)
von: Ranasinghe, Nishath Rajiv, et al.
Veröffentlicht: (2025)
The Transformative Influence of LLMs on Software Development & Developer Productivity
von: Jalil, Sajed
Veröffentlicht: (2023)
von: Jalil, Sajed
Veröffentlicht: (2023)
LLMLOOP: Improving LLM-Generated Code and Tests through Automated Iterative Feedback Loops
von: Ravi, Ravin, et al.
Veröffentlicht: (2026)
von: Ravi, Ravin, et al.
Veröffentlicht: (2026)
AcTracer: Active Testing of Large Language Model via Multi-Stage Sampling
von: Huang, Yuheng, et al.
Veröffentlicht: (2024)
von: Huang, Yuheng, et al.
Veröffentlicht: (2024)
Bug In the Code Stack: Can LLMs Find Bugs in Large Python Code Stacks
von: Lee, Hokyung, et al.
Veröffentlicht: (2024)
von: Lee, Hokyung, et al.
Veröffentlicht: (2024)
RMCBench: Benchmarking Large Language Models' Resistance to Malicious Code
von: Chen, Jiachi, et al.
Veröffentlicht: (2024)
von: Chen, Jiachi, et al.
Veröffentlicht: (2024)
Feedback-Normalized Developer Memory for Reinforcement-Learning Coding Agents: A Safety-Gated MCP Architecture
von: Iscan, Mehmet
Veröffentlicht: (2026)
von: Iscan, Mehmet
Veröffentlicht: (2026)
N-Version Assessment and Enhancement of Generative AI
von: Kessel, Marcus, et al.
Veröffentlicht: (2024)
von: Kessel, Marcus, et al.
Veröffentlicht: (2024)
Analyzing the Adoption of Database Management Systems Throughout the History of Open Source Projects
von: Paiva, Camila A., et al.
Veröffentlicht: (2026)
von: Paiva, Camila A., et al.
Veröffentlicht: (2026)
VulScribeR: Exploring RAG-based Vulnerability Augmentation with LLMs
von: Daneshvar, Seyed Shayan, et al.
Veröffentlicht: (2024)
von: Daneshvar, Seyed Shayan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Generative transformations and patterns in LLM-native approaches for software verification and falsification
von: Braberman, Víctor A., et al.
Veröffentlicht: (2024) -
Automated structural testing of LLM-based agents: methods, framework, and case studies
von: Kohl, Jens, et al.
Veröffentlicht: (2026) -
AgentEval: DAG-Structured Step-Level Evaluation for Agentic Workflows with Error Propagation Tracking
von: Guo, Dongxin, et al.
Veröffentlicht: (2026) -
AgentModernize: Preserving Business Logic in Legacy Modernization with Multi-Agent LLMs and Behavioral Specification Graphs
von: Ahmed, Sheikh Nazib, et al.
Veröffentlicht: (2026) -
Towards Observation Lakehouses: Living, Interactive Archives of Software Behavior
von: Kessel, Marcus
Veröffentlicht: (2025)