LLM4Vuln: A Unified Evaluation Framework for Decoupling and Enhancing LLMs' Vulnerability Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Sun, Yuqiang, Wu, Daoyuan, Xue, Yue, Liu, Han, Ma, Wei, Zhang, Lyuye, Liu, Yang, Li, Yingjiu |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ACFIX: Guiding LLMs with Mined Common RBAC Practices for Context-Aware Repair of Access Control Vulnerabilities in Smart Contracts
by: Zhang, Lyuye, et al.
Published: (2024)
by: Zhang, Lyuye, et al.
Published: (2024)
VulnRepairEval: An Exploit-Based Evaluation Framework for Assessing Large Language Model Vulnerability Repair Capabilities
by: Wang, Weizhe, et al.
Published: (2025)
by: Wang, Weizhe, et al.
Published: (2025)
GPTScan: Detecting Logic Vulnerabilities in Smart Contracts by Combining GPT with Program Analysis
by: Sun, Yuqiang, et al.
Published: (2023)
by: Sun, Yuqiang, et al.
Published: (2023)
Do Fine-Tuned LLMs Understand Vulnerabilities? An Investigation into the Semantic Trap
by: Huang, Feiyang, et al.
Published: (2026)
by: Huang, Feiyang, et al.
Published: (2026)
AutoVulnPHP: LLM-Powered Two-Stage PHP Vulnerability Detection and Automated Localization
by: Wang, Zhiqiang, et al.
Published: (2026)
by: Wang, Zhiqiang, et al.
Published: (2026)
Real-World Usability of Vulnerability Proof-of-Concepts: A Comprehensive Study
by: Dang, Wenjing, et al.
Published: (2025)
by: Dang, Wenjing, et al.
Published: (2025)
CrossCommitVuln-Bench: A Dataset of Multi-Commit Python Vulnerabilities Invisible to Per-Commit Static Analysis
by: Majumdar, Arunabh
Published: (2026)
by: Majumdar, Arunabh
Published: (2026)
MulVuln: Enhancing Pre-trained LMs with Shared and Language-Specific Knowledge for Multilingual Vulnerability Detection
by: Nguyen, Van, et al.
Published: (2025)
by: Nguyen, Van, et al.
Published: (2025)
Semantics-Aligned, Curriculum-Driven, and Reasoning-Enhanced Vulnerability Repair Framework
by: Yang, Chengran, et al.
Published: (2025)
by: Yang, Chengran, et al.
Published: (2025)
Rethinking and Exploring String-Based Malware Family Classification in the Era of LLMs and RAG
by: Chen, Yufan, et al.
Published: (2025)
by: Chen, Yufan, et al.
Published: (2025)
From LLMs to Agents: A Comparative Evaluation of LLMs and LLM-based Agents in Security Patch Detection
by: Han, Junxiao, et al.
Published: (2025)
by: Han, Junxiao, et al.
Published: (2025)
Evaluating Large Language Models for Line-Level Vulnerability Localization
by: Zhang, Jian, et al.
Published: (2024)
by: Zhang, Jian, et al.
Published: (2024)
Smart Contract Fuzzing Towards Profitable Vulnerabilities
by: Kong, Ziqiao, et al.
Published: (2025)
by: Kong, Ziqiao, et al.
Published: (2025)
VulInstruct: Teaching LLMs Root-Cause Reasoning for Vulnerability Detection via Security Specifications
by: Zhu, Hao, et al.
Published: (2025)
by: Zhu, Hao, et al.
Published: (2025)
PATCHEVAL: A New Benchmark for Evaluating LLMs on Patching Real-World Vulnerabilities
by: Wei, Zichao, et al.
Published: (2025)
by: Wei, Zichao, et al.
Published: (2025)
Execution-State-Aware LLM Reasoning for Automated Proof-of-Vulnerability Generation
by: Li, Haoyu, et al.
Published: (2026)
by: Li, Haoyu, et al.
Published: (2026)
How Code Representation Shapes False-Positive Dynamics in Cross-Language LLM Vulnerability Detection
by: Chen, Maofei, et al.
Published: (2026)
by: Chen, Maofei, et al.
Published: (2026)
An Empirical Study of Vulnerable Package Dependencies in LLM Repositories
by: Liu, Shuhan, et al.
Published: (2025)
by: Liu, Shuhan, et al.
Published: (2025)
LLM Agents for Automated Web Vulnerability Reproduction: Are We There Yet?
by: Liu, Bin, et al.
Published: (2025)
by: Liu, Bin, et al.
Published: (2025)
Towards Demystifying and Repairing LLM-in-the-Loop Vulnerabilities
by: Ma, Yujie, et al.
Published: (2026)
by: Ma, Yujie, et al.
Published: (2026)
Bridging Expert Reasoning and LLM Detection: A Knowledge-Driven Framework for Malicious Packages
by: Guo, Wenbo, et al.
Published: (2026)
by: Guo, Wenbo, et al.
Published: (2026)
LLM4CVE: Enabling Iterative Automated Vulnerability Repair with Large Language Models
by: Fakih, Mohamad, et al.
Published: (2025)
by: Fakih, Mohamad, et al.
Published: (2025)
KernJC: Automated Vulnerable Environment Generation for Linux Kernel Vulnerabilities
by: Ruan, Bonan, et al.
Published: (2024)
by: Ruan, Bonan, et al.
Published: (2024)
Exposing and Defending Membership Leakage in Vulnerability Prediction Models
by: Liao, Yihan, et al.
Published: (2025)
by: Liao, Yihan, et al.
Published: (2025)
Guiding Symbolic Execution with Static Analysis and LLMs for Vulnerability Discovery
by: Shafiuzzaman, Md, et al.
Published: (2026)
by: Shafiuzzaman, Md, et al.
Published: (2026)
To Err is Machine: Vulnerability Detection Challenges LLM Reasoning
by: Steenhoek, Benjamin, et al.
Published: (2024)
by: Steenhoek, Benjamin, et al.
Published: (2024)
Fast and Accurate Silent Vulnerability Fix Retrieval
by: Liu, Xueqing, et al.
Published: (2025)
by: Liu, Xueqing, et al.
Published: (2025)
Automated TEE Adaptation with LLMs: Identifying, Transforming, and Porting Sensitive Functions in Programs
by: Han, Ruidong, et al.
Published: (2025)
by: Han, Ruidong, et al.
Published: (2025)
VERCATION: Precise Vulnerable Open-source Software Version Identification based on Static Analysis and LLM
by: Cheng, Yiran, et al.
Published: (2024)
by: Cheng, Yiran, et al.
Published: (2024)
From Reviewers' Lens: Understanding Bug Bounty Report Invalid Reasons with LLMs
by: Zheng, Jiangrui, et al.
Published: (2025)
by: Zheng, Jiangrui, et al.
Published: (2025)
All Your Tokens are Belong to Us: Demystifying Address Verification Vulnerabilities in Solidity Smart Contracts
by: Sun, Tianle, et al.
Published: (2024)
by: Sun, Tianle, et al.
Published: (2024)
PatchSeeker: Mapping NVD Records to their Vulnerability-fixing Commits with LLM Generated Commits and Embeddings
by: Nguyen, Huu Hung, et al.
Published: (2025)
by: Nguyen, Huu Hung, et al.
Published: (2025)
CKGFuzzer: LLM-Based Fuzz Driver Generation Enhanced By Code Knowledge Graph
by: Xu, Hanxiang, et al.
Published: (2024)
by: Xu, Hanxiang, et al.
Published: (2024)
Coca: Improving and Explaining Graph Neural Network-Based Vulnerability Detection Systems
by: Cao, Sicong, et al.
Published: (2024)
by: Cao, Sicong, et al.
Published: (2024)
Finding Missing Input Validation in TEEs via LLM-Assisted Symbolic Execution
by: Ma, Chengyan, et al.
Published: (2026)
by: Ma, Chengyan, et al.
Published: (2026)
Casting a SPELL: Sentence Pairing Exploration for LLM Limitation-breaking
by: Huang, Yifan, et al.
Published: (2025)
by: Huang, Yifan, et al.
Published: (2025)
CleanVul: Automatic Function-Level Vulnerability Detection in Code Commits Using LLM Heuristics
by: Li, Yikun, et al.
Published: (2024)
by: Li, Yikun, et al.
Published: (2024)
Vulnerability Detection with Interprocedural Context in Multiple Languages: Assessing Effectiveness and Cost of Modern LLMs
by: Lira, Kevin, et al.
Published: (2026)
by: Lira, Kevin, et al.
Published: (2026)
VulZoo: A Comprehensive Vulnerability Intelligence Dataset
by: Ruan, Bonan, et al.
Published: (2024)
by: Ruan, Bonan, et al.
Published: (2024)
Evaluating LLMs for One-Shot Patching of Real and Artificial Vulnerabilities
by: Garg, Aayush, et al.
Published: (2025)
by: Garg, Aayush, et al.
Published: (2025)
Similar Items
-
ACFIX: Guiding LLMs with Mined Common RBAC Practices for Context-Aware Repair of Access Control Vulnerabilities in Smart Contracts
by: Zhang, Lyuye, et al.
Published: (2024) -
VulnRepairEval: An Exploit-Based Evaluation Framework for Assessing Large Language Model Vulnerability Repair Capabilities
by: Wang, Weizhe, et al.
Published: (2025) -
GPTScan: Detecting Logic Vulnerabilities in Smart Contracts by Combining GPT with Program Analysis
by: Sun, Yuqiang, et al.
Published: (2023) -
Do Fine-Tuned LLMs Understand Vulnerabilities? An Investigation into the Semantic Trap
by: Huang, Feiyang, et al.
Published: (2026) -
AutoVulnPHP: LLM-Powered Two-Stage PHP Vulnerability Detection and Automated Localization
by: Wang, Zhiqiang, et al.
Published: (2026)