From Program Slices to Causal Clarity: Evaluating Faithful, Actionable LLM-Generated Failure Explanations via Context Partitioning and LLM-as-a-Judge
Fuente:
arXiv
Saved in:
| Main Authors: | Porbeck, Julius, Adriano, Christian Medeiros, Giese, Holger |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LLM Based Input Space Partitioning Testing for Library APIs
by: Li, Jiageng, et al.
Published: (2024)
by: Li, Jiageng, et al.
Published: (2024)
From Feedback to Failure: Automated Android Performance Issue Reproduction
by: Li, Zhengquan, et al.
Published: (2025)
by: Li, Zhengquan, et al.
Published: (2025)
PALM: Path-aware LLM-based Test Generation with Comprehension
by: Wu, Yaoxuan, et al.
Published: (2025)
by: Wu, Yaoxuan, et al.
Published: (2025)
Flow-of-Action: SOP Enhanced LLM-Based Multi-Agent System for Root Cause Analysis
by: Pei, Changhua, et al.
Published: (2025)
by: Pei, Changhua, et al.
Published: (2025)
How Execution Features Relate to Failures: An Empirical Study and Diagnosis Approach
by: Smytzek, Marius, et al.
Published: (2025)
by: Smytzek, Marius, et al.
Published: (2025)
Mining Bug Repositories for Multi-Fault Programs
by: Callaghan, Dylan, et al.
Published: (2024)
by: Callaghan, Dylan, et al.
Published: (2024)
Automated Test Generation from Program Documentation Encoded in Code Comments
by: Denaro, Giovanni, et al.
Published: (2025)
by: Denaro, Giovanni, et al.
Published: (2025)
Analyzing Quantum Programs with LintQ: A Static Analysis Framework for Qiskit
by: Paltenghi, Matteo, et al.
Published: (2023)
by: Paltenghi, Matteo, et al.
Published: (2023)
CUJBench: Benchmarking LLM-Agent on Cross-Modal Failure Diagnosis from Browser to Backend
by: Meng, Haoming
Published: (2026)
by: Meng, Haoming
Published: (2026)
Synthesizing Performance Constraints for Evaluating and Improving Code Efficiency
by: Yang, Jun, et al.
Published: (2025)
by: Yang, Jun, et al.
Published: (2025)
A Catalog of Transformations to Remove Smells From Natural Language Tests
by: Aranda, Manoel, et al.
Published: (2024)
by: Aranda, Manoel, et al.
Published: (2024)
ETrace:Event-Driven Vulnerability Detection in Smart Contracts via LLM-Based Trace Analysis
by: Peng, Chenyang, et al.
Published: (2025)
by: Peng, Chenyang, et al.
Published: (2025)
Validating Formal Specifications with LLM-generated Test Cases
by: Cunha, Alcino, et al.
Published: (2025)
by: Cunha, Alcino, et al.
Published: (2025)
Comparing Human and LLM Generated Code: The Jury is Still Out!
by: Licorish, Sherlock A., et al.
Published: (2025)
by: Licorish, Sherlock A., et al.
Published: (2025)
Combined Program Analysis Techniques: A Systematic Mapping Study
by: Braione, Pietro, et al.
Published: (2026)
by: Braione, Pietro, et al.
Published: (2026)
Abstain and Validate: A Dual-LLM Policy for Reducing Noise in Agentic Program Repair
by: Cambronero, José, et al.
Published: (2025)
by: Cambronero, José, et al.
Published: (2025)
Auto-repair without test cases: How LLMs fix compilation errors in large industrial embedded code
by: Fu, Han, et al.
Published: (2025)
by: Fu, Han, et al.
Published: (2025)
Theory of Troubleshooting: The Developer's Cognitive Experience of Overcoming Confusion
by: Starr, Arty, et al.
Published: (2026)
by: Starr, Arty, et al.
Published: (2026)
Multi-modal Traffic Scenario Generation for Autonomous Driving System Testing
by: Tu, Zhi, et al.
Published: (2025)
by: Tu, Zhi, et al.
Published: (2025)
PathFuzzing: Worst Case Analysis by Fuzzing Symbolic-Execution Paths
by: Chen, Zimu, et al.
Published: (2025)
by: Chen, Zimu, et al.
Published: (2025)
Refining Fuzzed Crashing Inputs for Better Fault Diagnosis
by: Kim, Kieun, et al.
Published: (2025)
by: Kim, Kieun, et al.
Published: (2025)
ErrorPrism: Reconstructing Error Propagation Paths in Cloud Service Systems
by: Pu, Junsong, et al.
Published: (2025)
by: Pu, Junsong, et al.
Published: (2025)
In industrial embedded software, are some compilation errors easier to localize and fix than others?
by: Fu, Han, et al.
Published: (2024)
by: Fu, Han, et al.
Published: (2024)
How Do Developers Structure Unit Test Cases? An Empirical Study from the "AAA" Perspective
by: Wei, Chenhao, et al.
Published: (2024)
by: Wei, Chenhao, et al.
Published: (2024)
Towards debiasing code review support
by: Jetzen, Tobias, et al.
Published: (2024)
by: Jetzen, Tobias, et al.
Published: (2024)
CASCADE: Detecting Inconsistencies between Code and Documentation with Automatic Test Generation
by: Kiecker, Tobias, et al.
Published: (2026)
by: Kiecker, Tobias, et al.
Published: (2026)
VLM-Fuzz: Vision Language Model Assisted Recursive Depth-first Search Exploration for Effective UI Testing of Android Apps
by: Demissie, Biniam Fisseha, et al.
Published: (2025)
by: Demissie, Biniam Fisseha, et al.
Published: (2025)
SETBVE: Quality-Diversity Driven Exploration of Software Boundary Behaviors
by: Akbarova, Sabinakhon, et al.
Published: (2025)
by: Akbarova, Sabinakhon, et al.
Published: (2025)
PinChecker: Identifying Unsound Safe Abstractions of Rust Pinning APIs
by: Dai, Yuxuan, et al.
Published: (2025)
by: Dai, Yuxuan, et al.
Published: (2025)
Empirical Analysis of Temporal and Spatial Fault Characteristics in Multi-Fault Bug Repositories
by: Callaghan, Dylan, et al.
Published: (2025)
by: Callaghan, Dylan, et al.
Published: (2025)
Plug it and Play on Logs: A Configuration-Free Statistic-Based Log Parser
by: Qin, Qiaolin, et al.
Published: (2025)
by: Qin, Qiaolin, et al.
Published: (2025)
A Comprehensive Study on Large Language Models for Mutation Testing
by: Wang, Bo, et al.
Published: (2024)
by: Wang, Bo, et al.
Published: (2024)
E-Test: E'er-Improving Test Suites
by: Qiu, Ketai, et al.
Published: (2025)
by: Qiu, Ketai, et al.
Published: (2025)
Automated Generation of Issue-Reproducing Tests by Combining LLMs and Search-Based Testing
by: Kitsios, Konstantinos, et al.
Published: (2025)
by: Kitsios, Konstantinos, et al.
Published: (2025)
Mutation Analysis with Execution Taints
by: Gopinath, Rahul, et al.
Published: (2024)
by: Gopinath, Rahul, et al.
Published: (2024)
Which Combination of Test Metrics Can Predict Success of a Software Project? A Case Study in a Year-Long Project Course
by: Filipovic, Marina, et al.
Published: (2024)
by: Filipovic, Marina, et al.
Published: (2024)
On Interaction Effects in Greybox Fuzzing
by: Kitsios, Konstantinos, et al.
Published: (2025)
by: Kitsios, Konstantinos, et al.
Published: (2025)
Leveraging Propagated Infection to Crossfire Mutants
by: Du, Hang, et al.
Published: (2024)
by: Du, Hang, et al.
Published: (2024)
Utilizing Precise and Complete Code Context to Guide LLM in Automatic False Positive Mitigation
by: Chen, Jinbao, et al.
Published: (2024)
by: Chen, Jinbao, et al.
Published: (2024)
Mining Subscenario Refactoring Opportunities in Behaviour-Driven Software Test Suites: ML Classifiers and LLM-Judge Baselines
by: Mughal, Ali Hassaan, et al.
Published: (2026)
by: Mughal, Ali Hassaan, et al.
Published: (2026)
Similar Items
-
LLM Based Input Space Partitioning Testing for Library APIs
by: Li, Jiageng, et al.
Published: (2024) -
From Feedback to Failure: Automated Android Performance Issue Reproduction
by: Li, Zhengquan, et al.
Published: (2025) -
PALM: Path-aware LLM-based Test Generation with Comprehension
by: Wu, Yaoxuan, et al.
Published: (2025) -
Flow-of-Action: SOP Enhanced LLM-Based Multi-Agent System for Root Cause Analysis
by: Pei, Changhua, et al.
Published: (2025) -
How Execution Features Relate to Failures: An Empirical Study and Diagnosis Approach
by: Smytzek, Marius, et al.
Published: (2025)