FaultLine: Automated Proof-of-Vulnerability Generation Using LLM Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Nitin, Vikram, Ray, Baishakhi, Moghaddam, Roshanak Zilouchian |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PerfBench: Can Agents Resolve Real-World Performance Bugs?
by: Garg, Spandan, et al.
Published: (2025)
by: Garg, Spandan, et al.
Published: (2025)
SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments
by: Arora, Avi, et al.
Published: (2025)
by: Arora, Avi, et al.
Published: (2025)
SpecTra: Enhancing the Code Translation Ability of Language Models by Generating Multi-Modal Specifications
by: Nitin, Vikram, et al.
Published: (2024)
by: Nitin, Vikram, et al.
Published: (2024)
The SWE-Bench Illusion: When State-of-the-Art LLMs Remember Instead of Reason
by: Liang, Shanchao, et al.
Published: (2025)
by: Liang, Shanchao, et al.
Published: (2025)
RAPGen: An Approach for Fixing Code Inefficiencies in Zero-Shot
by: Garg, Spandan, et al.
Published: (2023)
by: Garg, Spandan, et al.
Published: (2023)
AutoDev: Automated AI-Driven Development
by: Tufano, Michele, et al.
Published: (2024)
by: Tufano, Michele, et al.
Published: (2024)
Yuga: Automatically Detecting Lifetime Annotation Bugs in the Rust Language
by: Nitin, Vikram, et al.
Published: (2023)
by: Nitin, Vikram, et al.
Published: (2023)
Closing the Gap: A User Study on the Real-world Usefulness of AI-powered Vulnerability Detection & Repair in the IDE
by: Steenhoek, Benjamin, et al.
Published: (2024)
by: Steenhoek, Benjamin, et al.
Published: (2024)
C2SaferRust: Transforming C Projects into Safer Rust with NeuroSymbolic Techniques
by: Nitin, Vikram, et al.
Published: (2025)
by: Nitin, Vikram, et al.
Published: (2025)
RefactorBench: Evaluating Stateful Reasoning in Language Agents Through Code
by: Gautam, Dhruv, et al.
Published: (2025)
by: Gautam, Dhruv, et al.
Published: (2025)
Red Teaming Program Repair Agents: When Correct Patches can Hide Vulnerabilities
by: Chen, Simin, et al.
Published: (2025)
by: Chen, Simin, et al.
Published: (2025)
Code Quality Analysis of Translations from C to Rust
by: Tadesse, Biruk, et al.
Published: (2026)
by: Tadesse, Biruk, et al.
Published: (2026)
Copilot Evaluation Harness: Evaluating LLM-Guided Software Programming
by: Agarwal, Anisha, et al.
Published: (2024)
by: Agarwal, Anisha, et al.
Published: (2024)
Execution-State-Aware LLM Reasoning for Automated Proof-of-Vulnerability Generation
by: Li, Haoyu, et al.
Published: (2026)
by: Li, Haoyu, et al.
Published: (2026)
Automated Code Editing with Search-Generate-Modify
by: Liu, Changshu, et al.
Published: (2023)
by: Liu, Changshu, et al.
Published: (2023)
Understanding Automated Program Repair Agents Through the Lens of Traceability: An Empirical Study
by: Ceka, Ira, et al.
Published: (2025)
by: Ceka, Ira, et al.
Published: (2025)
SemAgent: A Semantics Aware Program Repair Agent
by: Pabba, Anvith, et al.
Published: (2025)
by: Pabba, Anvith, et al.
Published: (2025)
Can LLM Prompting Serve as a Proxy for Static Analysis in Vulnerability Detection
by: Ceka, Ira, et al.
Published: (2024)
by: Ceka, Ira, et al.
Published: (2024)
Trustworthy AI Software Engineers
by: Aleti, Aldeida, et al.
Published: (2026)
by: Aleti, Aldeida, et al.
Published: (2026)
Program Analysis Guided LLM Agent for Proof-of-Concept Generation
by: Desai, Achintya, et al.
Published: (2026)
by: Desai, Achintya, et al.
Published: (2026)
CWEval: Outcome-driven Evaluation on Functionality and Security of LLM Code Generation
by: Peng, Jinjun, et al.
Published: (2025)
by: Peng, Jinjun, et al.
Published: (2025)
LLM Agents for Automated Web Vulnerability Reproduction: Are We There Yet?
by: Liu, Bin, et al.
Published: (2025)
by: Liu, Bin, et al.
Published: (2025)
A Systematic Study on Generating Web Vulnerability Proof-of-Concepts Using Large Language Models
by: Zhao, Mengyao, et al.
Published: (2025)
by: Zhao, Mengyao, et al.
Published: (2025)
REFINE: Enhancing Program Repair Agents through Context-Aware Patch Refinement
by: Pabba, Anvith, et al.
Published: (2025)
by: Pabba, Anvith, et al.
Published: (2025)
CYCLE: Learning to Self-Refine the Code Generation
by: Ding, Yangruibo, et al.
Published: (2024)
by: Ding, Yangruibo, et al.
Published: (2024)
VulnResolver: A Hybrid Agent Framework for LLM-Based Automated Vulnerability Issue Resolution
by: Zhang, Mingming, et al.
Published: (2026)
by: Zhang, Mingming, et al.
Published: (2026)
Fuzzing with Agents? Generators Are All You Need
by: Vikram, Vasudev, et al.
Published: (2026)
by: Vikram, Vasudev, et al.
Published: (2026)
UTFix: Change Aware Unit Test Repairing using LLM
by: Rahman, Shanto, et al.
Published: (2025)
by: Rahman, Shanto, et al.
Published: (2025)
Vulnerability Detection with Code Language Models: How Far Are We?
by: Ding, Yangruibo, et al.
Published: (2024)
by: Ding, Yangruibo, et al.
Published: (2024)
A Dual-Loop Agent Framework for Automated Vulnerability Reproduction
by: Liu, Bin, et al.
Published: (2026)
by: Liu, Bin, et al.
Published: (2026)
Code-Aware Prompting: A study of Coverage Guided Test Generation in Regression Setting using LLM
by: Ryan, Gabriel, et al.
Published: (2024)
by: Ryan, Gabriel, et al.
Published: (2024)
Towards Causal Deep Learning for Vulnerability Detection
by: Rahman, Md Mahbubur, et al.
Published: (2023)
by: Rahman, Md Mahbubur, et al.
Published: (2023)
Agentic AI Software Engineers: Programming with Trust
by: Roychoudhury, Abhik, et al.
Published: (2025)
by: Roychoudhury, Abhik, et al.
Published: (2025)
Dynamic Benchmarking of Reasoning Capabilities in Code Large Language Models Under Data Contamination
by: Chen, Simin, et al.
Published: (2025)
by: Chen, Simin, et al.
Published: (2025)
Your Compiler is Backdooring Your Model: Understanding and Exploiting Compilation Inconsistency Vulnerabilities in Deep Learning Compilers
by: Chen, Simin, et al.
Published: (2025)
by: Chen, Simin, et al.
Published: (2025)
LLMs in Code Vulnerability Analysis: A Proof of Concept
by: Sultana, Shaznin, et al.
Published: (2026)
by: Sultana, Shaznin, et al.
Published: (2026)
LLM Agents for Automated Dependency Upgrades
by: Tawosi, Vali, et al.
Published: (2025)
by: Tawosi, Vali, et al.
Published: (2025)
From Trace to Line: LLM Agent for Real-World OSS Vulnerability Localization
by: Xi, Haoran, et al.
Published: (2025)
by: Xi, Haoran, et al.
Published: (2025)
Outrunning LLM Cutoffs: A Live Kernel Crash Resolution Benchmark for All
by: Huang, Chenxi, et al.
Published: (2026)
by: Huang, Chenxi, et al.
Published: (2026)
Terminus-4B: Can a Smaller Model Replace Frontier LLMs at Agentic Execution Tasks?
by: Garg, Spandan, et al.
Published: (2026)
by: Garg, Spandan, et al.
Published: (2026)
Similar Items
-
PerfBench: Can Agents Resolve Real-World Performance Bugs?
by: Garg, Spandan, et al.
Published: (2025) -
SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments
by: Arora, Avi, et al.
Published: (2025) -
SpecTra: Enhancing the Code Translation Ability of Language Models by Generating Multi-Modal Specifications
by: Nitin, Vikram, et al.
Published: (2024) -
The SWE-Bench Illusion: When State-of-the-Art LLMs Remember Instead of Reason
by: Liang, Shanchao, et al.
Published: (2025) -
RAPGen: An Approach for Fixing Code Inefficiencies in Zero-Shot
by: Garg, Spandan, et al.
Published: (2023)