AgentSZZ: Teaching the LLM Agent to Play Detective with Bug-Inducing Commits
Fuente:
arXiv
Saved in:
| Main Authors: | Lyu, Yunbo, Shi, Jieke, Kang, Hong Jin, Widyasari, Ratnadira, He, Junda, Niu, Yuqing, Yang, Chengran, Chen, Junkai, Yang, Zhou, Lawall, Julia, Lo, David |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Evaluating SZZ Implementations: An Empirical Study on the Linux Kernel
by: Lyu, Yunbo, et al.
Published: (2023)
by: Lyu, Yunbo, et al.
Published: (2023)
Beyond ChatGPT: Enhancing Software Quality Assurance Tasks with Diverse LLMs and Validation Techniques
by: Widyasari, Ratnadira, et al.
Published: (2024)
by: Widyasari, Ratnadira, et al.
Published: (2024)
Back to the Basics: Rethinking Issue-Commit Linking with LLM-Assisted Retrieval
by: Huang, Huihui, et al.
Published: (2025)
by: Huang, Huihui, et al.
Published: (2025)
AgenticSZZ: Temporal Knowledge Graph-Guided Agentic Bug-Inducing Commit Identification
by: Shi, Yu, et al.
Published: (2026)
by: Shi, Yu, et al.
Published: (2026)
Confident Learning-based Network for Detecting Bug-Inducing Commits on SZZ with Noisy Labels
by: Sun, Weihao, et al.
Published: (2026)
by: Sun, Weihao, et al.
Published: (2026)
MAS-SZZ: Multi-Agentic SZZ Algorithm for Vulnerability-Inducing Commit Identification
by: Cao, Sicong, et al.
Published: (2026)
by: Cao, Sicong, et al.
Published: (2026)
SLICEMATE: Accurate and Scalable Static Program Slicing via LLM-Powered Agents
by: Chang, Jianming, et al.
Published: (2025)
by: Chang, Jianming, et al.
Published: (2025)
ACECode: A Reinforcement Learning Framework for Aligning Code Efficiency and Correctness in Code Language Models
by: Yang, Chengran, et al.
Published: (2024)
by: Yang, Chengran, et al.
Published: (2024)
Debt Behind the AI Boom: A Large-Scale Empirical Study of AI-Generated Code in the Wild
by: Liu, Yue, et al.
Published: (2026)
by: Liu, Yue, et al.
Published: (2026)
PenForge: On-the-Fly Expert Agent Construction for Automated Penetration Testing
by: Huang, Huihui, et al.
Published: (2026)
by: Huang, Huihui, et al.
Published: (2026)
LessLeak-Bench: A First Investigation of Data Leakage in LLMs Across 83 Software Engineering Benchmarks
by: Zhou, Xin, et al.
Published: (2025)
by: Zhou, Xin, et al.
Published: (2025)
Demystifying Faulty Code with LLM: Step-by-Step Reasoning for Explainable Fault Localization
by: Widyasari, Ratnadira, et al.
Published: (2024)
by: Widyasari, Ratnadira, et al.
Published: (2024)
SecureVibeBench: Benchmarking Secure Vibe Coding of AI Agents via Reconstructing Vulnerability-Introducing Scenarios
by: Chen, Junkai, et al.
Published: (2025)
by: Chen, Junkai, et al.
Published: (2025)
Think Like Human Developers: Harnessing Community Knowledge for Structured Code Reasoning
by: Yang, Chengran, et al.
Published: (2025)
by: Yang, Chengran, et al.
Published: (2025)
What You Trust Is Insecure: Demystifying How Developers (Mis)Use Trusted Execution Environments in Practice
by: Niu, Yuqing, et al.
Published: (2025)
by: Niu, Yuqing, et al.
Published: (2025)
Explaining Explanation: An Empirical Study on Explanation in Code Reviews
by: Widyasari, Ratnadira, et al.
Published: (2023)
by: Widyasari, Ratnadira, et al.
Published: (2023)
APIDocBooster: An Extract-Then-Abstract Framework Leveraging Large Language Models for Augmenting API Documentation
by: Yang, Chengran, et al.
Published: (2023)
by: Yang, Chengran, et al.
Published: (2023)
Curiosity-Driven Testing for Sequential Decision-Making Process
by: He, Junda, et al.
Published: (2025)
by: He, Junda, et al.
Published: (2025)
A Functional Software Reference Architecture for LLM-Integrated Systems
by: Bucaioni, Alessio, et al.
Published: (2025)
by: Bucaioni, Alessio, et al.
Published: (2025)
Mut4All: Fuzzing Compilers via LLM-Synthesized Mutators Learned from Bug Reports
by: Wang, Bo, et al.
Published: (2025)
by: Wang, Bo, et al.
Published: (2025)
Let the Trial Begin: A Mock-Court Approach to Vulnerability Detection using LLM-Based Agents
by: Widyasari, Ratnadira, et al.
Published: (2025)
by: Widyasari, Ratnadira, et al.
Published: (2025)
Compiling Code LLMs into Lightweight Executables
by: Shi, Jieke, et al.
Published: (2026)
by: Shi, Jieke, et al.
Published: (2026)
LLM4SZZ: Enhancing SZZ Algorithm with Context-Enhanced Assessment on Large Language Models
by: Tang, Lingxiao, et al.
Published: (2025)
by: Tang, Lingxiao, et al.
Published: (2025)
Do Existing Testing Tools Really Uncover Gender Bias in Text-to-Image Models?
by: Lyu, Yunbo, et al.
Published: (2025)
by: Lyu, Yunbo, et al.
Published: (2025)
Synthesizing Efficient and Permissive Programmatic Runtime Shields for Neural Policies
by: Shi, Jieke, et al.
Published: (2024)
by: Shi, Jieke, et al.
Published: (2024)
Greening Large Language Models of Code
by: Shi, Jieke, et al.
Published: (2023)
by: Shi, Jieke, et al.
Published: (2023)
Mapping NVD Records to Their Vulnerability-fixing Commits: How Hard is It?
by: Nguyen, Huu Hung, et al.
Published: (2025)
by: Nguyen, Huu Hung, et al.
Published: (2025)
LLM-Based Multi-Agent Systems for Software Engineering: Literature Review, Vision and the Road Ahead
by: He, Junda, et al.
Published: (2024)
by: He, Junda, et al.
Published: (2024)
Artificial Intelligence for Software Architecture: Literature Review and the Road Ahead
by: Bucaioni, Alessio, et al.
Published: (2025)
by: Bucaioni, Alessio, et al.
Published: (2025)
BugsInPy: A Database of Existing Bugs in Python Programs to Enable Controlled Testing and Debugging Studies
by: Widyasari, Ratnadira, et al.
Published: (2024)
by: Widyasari, Ratnadira, et al.
Published: (2024)
Finding Missing Input Validation in TEEs via LLM-Assisted Symbolic Execution
by: Ma, Chengyan, et al.
Published: (2026)
by: Ma, Chengyan, et al.
Published: (2026)
WIA-SZZ: Work Item Aware SZZ
by: Perez-Rosero, Salomé, et al.
Published: (2024)
by: Perez-Rosero, Salomé, et al.
Published: (2024)
"My productivity is boosted, but ..." Demystifying Users' Perception on AI Coding Assistants
by: Lyu, Yunbo, et al.
Published: (2025)
by: Lyu, Yunbo, et al.
Published: (2025)
How and Why Agents Can Identify Bug-Introducing Commits
by: Risse, Niklas, et al.
Published: (2026)
by: Risse, Niklas, et al.
Published: (2026)
An Empirical Study of Bugs in Modern LLM Agent Frameworks
by: Zhu, Xinxue, et al.
Published: (2026)
by: Zhu, Xinxue, et al.
Published: (2026)
Automated Repair of TEE Partitioning Issues via DSL-Guided and LLM-Assisted Patching
by: Ma, Chengyan, et al.
Published: (2026)
by: Ma, Chengyan, et al.
Published: (2026)
Efficient and Green Large Language Models for Software Engineering: Literature Review, Vision, and the Road Ahead
by: Shi, Jieke, et al.
Published: (2024)
by: Shi, Jieke, et al.
Published: (2024)
CleanVul: Automatic Function-Level Vulnerability Detection in Code Commits Using LLM Heuristics
by: Li, Yikun, et al.
Published: (2024)
by: Li, Yikun, et al.
Published: (2024)
Finding Safety Violations of AI-Enabled Control Systems through the Lens of Synthesized Proxy Programs
by: Shi, Jieke, et al.
Published: (2024)
by: Shi, Jieke, et al.
Published: (2024)
LLM-as-a-Judge for Software Engineering: Literature Review, Vision, and the Road Ahead
by: He, Junda, et al.
Published: (2025)
by: He, Junda, et al.
Published: (2025)
Similar Items
-
Evaluating SZZ Implementations: An Empirical Study on the Linux Kernel
by: Lyu, Yunbo, et al.
Published: (2023) -
Beyond ChatGPT: Enhancing Software Quality Assurance Tasks with Diverse LLMs and Validation Techniques
by: Widyasari, Ratnadira, et al.
Published: (2024) -
Back to the Basics: Rethinking Issue-Commit Linking with LLM-Assisted Retrieval
by: Huang, Huihui, et al.
Published: (2025) -
AgenticSZZ: Temporal Knowledge Graph-Guided Agentic Bug-Inducing Commit Identification
by: Shi, Yu, et al.
Published: (2026) -
Confident Learning-based Network for Detecting Bug-Inducing Commits on SZZ with Noisy Labels
by: Sun, Weihao, et al.
Published: (2026)