Does Pass Rate Tell the Whole Story? Evaluating Design Constraint Compliance in LLM-based Issue Resolution
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yu, Kai, Zhou, Zhenhao, Zeng, Junhao, Wang, Ying, Du, Xueying, Yuan, Zhiqiang, Liu, Junwei, Zhou, Ziyu, Wang, Yujia, Wang, Chong, Peng, Xin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Do Papers Tell the Whole Story? A Benchmark and Framework for Uncovering Hidden Implementation Gaps in Bioinformatics
von: Xu, Tianxiang, et al.
Veröffentlicht: (2026)
von: Xu, Tianxiang, et al.
Veröffentlicht: (2026)
Extracting Conceptual Knowledge to Locate Software Issues
von: Wang, Ying, et al.
Veröffentlicht: (2025)
von: Wang, Ying, et al.
Veröffentlicht: (2025)
Minimizing False Positives in Static Bug Detection via LLM-Enhanced Path Feasibility Analysis
von: Du, Xueying, et al.
Veröffentlicht: (2025)
von: Du, Xueying, et al.
Veröffentlicht: (2025)
LLM4Perf: Large Language Models Are Effective Samplers for Multi-Objective Performance Modeling
von: Wang, Xin, et al.
Veröffentlicht: (2025)
von: Wang, Xin, et al.
Veröffentlicht: (2025)
Compact Constraint Encoding for LLM Code Generation: An Empirical Study of Token Economics and Constraint Compliance
von: Tang, Hanzhang
Veröffentlicht: (2026)
von: Tang, Hanzhang
Veröffentlicht: (2026)
Beyond Pass/Fail: The Story of Learning-Based Testing
von: Rahman, Sheikh Md. Mushfiqur, et al.
Veröffentlicht: (2025)
von: Rahman, Sheikh Md. Mushfiqur, et al.
Veröffentlicht: (2025)
Taming System Complexity: Demystifying Software Engineering Agents in Diagnosing Linux Kernel Faults
von: Zhou, Zhenhao, et al.
Veröffentlicht: (2025)
von: Zhou, Zhenhao, et al.
Veröffentlicht: (2025)
Vul-RAG: Enhancing LLM-based Vulnerability Detection via Knowledge-level RAG
von: Du, Xueying, et al.
Veröffentlicht: (2024)
von: Du, Xueying, et al.
Veröffentlicht: (2024)
MCTS-Refined CoT: High-Quality Fine-Tuning Data for LLM-Based Repository Issue Resolution
von: Wang, Yibo, et al.
Veröffentlicht: (2025)
von: Wang, Yibo, et al.
Veröffentlicht: (2025)
Defects4Log: Benchmarking LLMs for Logging Code Defect Detection and Reasoning
von: Wang, Xin, et al.
Veröffentlicht: (2025)
von: Wang, Xin, et al.
Veröffentlicht: (2025)
OptiLoop: Coordination-in-the-Loop Verification and Repair for LLM-Generated Optimization Agents
von: Xu, Yujia, et al.
Veröffentlicht: (2026)
von: Xu, Yujia, et al.
Veröffentlicht: (2026)
MAGIS: LLM-Based Multi-Agent Framework for GitHub Issue Resolution
von: Tao, Wei, et al.
Veröffentlicht: (2024)
von: Tao, Wei, et al.
Veröffentlicht: (2024)
Enhancing LLM-based Fault Localization with a Functionality-Aware Retrieval-Augmented Generation Framework
von: Shi, Xinyu, et al.
Veröffentlicht: (2025)
von: Shi, Xinyu, et al.
Veröffentlicht: (2025)
Beyond Fixed Tests: Repository-Level Issue Resolution as Coevolution of Code and Behavioral Constraints
von: Li, Kefan, et al.
Veröffentlicht: (2026)
von: Li, Kefan, et al.
Veröffentlicht: (2026)
LLM-Based Test Case Generation in DBMS through Monte Carlo Tree Search
von: Chen, Yujia, et al.
Veröffentlicht: (2026)
von: Chen, Yujia, et al.
Veröffentlicht: (2026)
ThinkRepair: Self-Directed Automated Program Repair
von: Yin, Xin, et al.
Veröffentlicht: (2024)
von: Yin, Xin, et al.
Veröffentlicht: (2024)
LiCoEval: Evaluating LLMs on License Compliance in Code Generation
von: Xu, Weiwei, et al.
Veröffentlicht: (2024)
von: Xu, Weiwei, et al.
Veröffentlicht: (2024)
One-for-All Does Not Work! Enhancing Vulnerability Detection by Mixture-of-Experts (MoE)
von: Yang, Xu, et al.
Veröffentlicht: (2025)
von: Yang, Xu, et al.
Veröffentlicht: (2025)
MCBA: A Matroid Constraint-Based Approach for Composite Service Recommendation Considering Compatibility and Diversity
von: Sun, Ying, et al.
Veröffentlicht: (2024)
von: Sun, Ying, et al.
Veröffentlicht: (2024)
Position Paper: Programming Language Techniques for Bridging LLM Code Generation Semantic Gaps
von: Du, Yalong, et al.
Veröffentlicht: (2025)
von: Du, Yalong, et al.
Veröffentlicht: (2025)
EvoDev: An Iterative Feature-Driven Framework for End-to-End Software Development with LLM-based Agents
von: Liu, Junwei, et al.
Veröffentlicht: (2025)
von: Liu, Junwei, et al.
Veröffentlicht: (2025)
The Tool-Overuse Illusion: Why Does LLM Prefer External Tools over Internal Knowledge?
von: Zeng, Yirong, et al.
Veröffentlicht: (2026)
von: Zeng, Yirong, et al.
Veröffentlicht: (2026)
Predicting Issue Resolution Time of OSS Using Multiple Features
von: Yu Qiao, et al.
Veröffentlicht: (2024)
von: Yu Qiao, et al.
Veröffentlicht: (2024)
Code Copycat Conundrum: Demystifying Repetition in LLM-based Code Generation
von: Liu, Mingwei, et al.
Veröffentlicht: (2025)
von: Liu, Mingwei, et al.
Veröffentlicht: (2025)
How LLMs Aid in UML Modeling: An Exploratory Study with Novice Analysts
von: Wang, Beian, et al.
Veröffentlicht: (2024)
von: Wang, Beian, et al.
Veröffentlicht: (2024)
Breaking the Dependency Chaos: A Constraint-Driven Python Dependency Resolution Strategy with Selective LLM Imputation
von: Chowdhury, Kowshik, et al.
Veröffentlicht: (2026)
von: Chowdhury, Kowshik, et al.
Veröffentlicht: (2026)
PredicateFix: Repairing Static Analysis Alerts with Bridging Predicates
von: Xiao, Yuan-An, et al.
Veröffentlicht: (2025)
von: Xiao, Yuan-An, et al.
Veröffentlicht: (2025)
REAgent: Requirement-Driven LLM Agents for Software Issue Resolution
von: Kuang, Shiqi, et al.
Veröffentlicht: (2026)
von: Kuang, Shiqi, et al.
Veröffentlicht: (2026)
Evaluating and Improving Automated Repository-Level Rust Issue Resolution with LLM-based Agents
von: Xiang, Jiahong, et al.
Veröffentlicht: (2026)
von: Xiang, Jiahong, et al.
Veröffentlicht: (2026)
Mage: Multi-Axis Evaluation of LLM-Generated Executable Game Scenes Beyond Compile-Pass Rate
von: Liu, Hugh Xuechen, et al.
Veröffentlicht: (2026)
von: Liu, Hugh Xuechen, et al.
Veröffentlicht: (2026)
Debloating Software Through Enhanced Static Analysis and Constraint Rules
von: Xiaohu Song, et al.
Veröffentlicht: (2025)
von: Xiaohu Song, et al.
Veröffentlicht: (2025)
Mut4All: Fuzzing Compilers via LLM-Synthesized Mutators Learned from Bug Reports
von: Wang, Bo, et al.
Veröffentlicht: (2025)
von: Wang, Bo, et al.
Veröffentlicht: (2025)
Leveraging Large Vision Language Model For Better Automatic Web GUI Testing
von: Wang, Siyi, et al.
Veröffentlicht: (2024)
von: Wang, Siyi, et al.
Veröffentlicht: (2024)
Boosting Static Resource Leak Detection via LLM-based Resource-Oriented Intention Inference
von: Wang, Chong, et al.
Veröffentlicht: (2023)
von: Wang, Chong, et al.
Veröffentlicht: (2023)
Assessing Privacy Compliance of Android Third-Party SDKs
von: Meng, Mark Huasong, et al.
Veröffentlicht: (2024)
von: Meng, Mark Huasong, et al.
Veröffentlicht: (2024)
SWE-Mirror: Scaling Issue-Resolving Datasets by Mirroring Issues Across Repositories
von: Wang, Junhao, et al.
Veröffentlicht: (2025)
von: Wang, Junhao, et al.
Veröffentlicht: (2025)
AssertFlip: Reproducing Bugs via Inversion of LLM-Generated Passing Tests
von: Khatib, Lara, et al.
Veröffentlicht: (2025)
von: Khatib, Lara, et al.
Veröffentlicht: (2025)
CircuChain: Disentangling Competence and Compliance in LLM Circuit Analysis
von: Ravishankara, Mayank
Veröffentlicht: (2026)
von: Ravishankara, Mayank
Veröffentlicht: (2026)
VulnResolver: A Hybrid Agent Framework for LLM-Based Automated Vulnerability Issue Resolution
von: Zhang, Mingming, et al.
Veröffentlicht: (2026)
von: Zhang, Mingming, et al.
Veröffentlicht: (2026)
Code Digital Twin: Empowering LLMs with Tacit Knowledge for Complex Software Development
von: Peng, Xin, et al.
Veröffentlicht: (2025)
von: Peng, Xin, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Do Papers Tell the Whole Story? A Benchmark and Framework for Uncovering Hidden Implementation Gaps in Bioinformatics
von: Xu, Tianxiang, et al.
Veröffentlicht: (2026) -
Extracting Conceptual Knowledge to Locate Software Issues
von: Wang, Ying, et al.
Veröffentlicht: (2025) -
Minimizing False Positives in Static Bug Detection via LLM-Enhanced Path Feasibility Analysis
von: Du, Xueying, et al.
Veröffentlicht: (2025) -
LLM4Perf: Large Language Models Are Effective Samplers for Multi-Objective Performance Modeling
von: Wang, Xin, et al.
Veröffentlicht: (2025) -
Compact Constraint Encoding for LLM Code Generation: An Empirical Study of Token Economics and Constraint Compliance
von: Tang, Hanzhang
Veröffentlicht: (2026)