Beyond Single Bugs: Benchmarking Large Language Models for Multi-Vulnerability Detection
Fuente:
arXiv
Saved in:
| Main Authors: | Pushkar, Chinmay, Kabra, Sanchit, Kumar, Dhruv, Challa, Jagat Sesh |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
BugWhisperer: Fine-Tuning LLMs for SoC Hardware Vulnerability Detection
by: Tarek, Shams, et al.
Published: (2025)
by: Tarek, Shams, et al.
Published: (2025)
Beyond Content Safety: Real-Time Monitoring for Reasoning Vulnerabilities in Large Language Models
by: Wang, Xunguang, et al.
Published: (2026)
by: Wang, Xunguang, et al.
Published: (2026)
Enhancing Reverse Engineering: Investigating and Benchmarking Large Language Models for Vulnerability Analysis in Decompiled Binaries
by: Manuel, Dylan, et al.
Published: (2024)
by: Manuel, Dylan, et al.
Published: (2024)
Harnessing Large Language Models for Software Vulnerability Detection: A Comprehensive Benchmarking Study
by: Tamberg, Karl, et al.
Published: (2024)
by: Tamberg, Karl, et al.
Published: (2024)
MARVEL: Multi-Agent RTL Vulnerability Extraction using Large Language Models
by: Collini, Luca, et al.
Published: (2025)
by: Collini, Luca, et al.
Published: (2025)
Finetuning Large Language Models for Vulnerability Detection
by: Shestov, Alexey, et al.
Published: (2024)
by: Shestov, Alexey, et al.
Published: (2024)
CryptoScope: Utilizing Large Language Models for Automated Cryptographic Logic Vulnerability Detection
by: Li, Zhihao, et al.
Published: (2025)
by: Li, Zhihao, et al.
Published: (2025)
Retrieval Augmented Generation Integrated Large Language Models in Smart Contract Vulnerability Detection
by: Yu, Jeffy
Published: (2024)
by: Yu, Jeffy
Published: (2024)
Transient Turn Injection: Exposing Stateless Multi-Turn Vulnerabilities in Large Language Models
by: Rayhan, Naheed, et al.
Published: (2026)
by: Rayhan, Naheed, et al.
Published: (2026)
CloudLens: Modeling and Detecting Cloud Security Vulnerabilities
by: Kazdagli, Mikhail, et al.
Published: (2024)
by: Kazdagli, Mikhail, et al.
Published: (2024)
Evaluating Adversarial Vulnerabilities in Modern Large Language Models
by: Perel, Tom
Published: (2025)
by: Perel, Tom
Published: (2025)
LLbezpeky: Leveraging Large Language Models for Vulnerability Detection
by: Mathews, Noble Saji, et al.
Published: (2024)
by: Mathews, Noble Saji, et al.
Published: (2024)
Breaking the Loop: Detecting and Mitigating Denial-of-Service Vulnerabilities in Large Language Models
by: Yu, Junzhe, et al.
Published: (2025)
by: Yu, Junzhe, et al.
Published: (2025)
Exploring Membership Inference Vulnerabilities in Clinical Large Language Models
by: Nemecek, Alexander, et al.
Published: (2025)
by: Nemecek, Alexander, et al.
Published: (2025)
ParaVul: A Parallel Large Language Model and Retrieval-Augmented Framework for Smart Contract Vulnerability Detection
by: Huang, Tenghui, et al.
Published: (2025)
by: Huang, Tenghui, et al.
Published: (2025)
Unvalidated Trust: Cross-Stage Vulnerabilities in Large Language Model Architectures
by: Schwarz, Dominik
Published: (2025)
by: Schwarz, Dominik
Published: (2025)
Beyond Detection: Leveraging Large Language Models for Cyber Attack Prediction in IoT Networks
by: Diaf, Alaeddine, et al.
Published: (2024)
by: Diaf, Alaeddine, et al.
Published: (2024)
Evaluating Large Language Models for Security Bug Report Prediction
by: Soltaniani, Farnaz, et al.
Published: (2026)
by: Soltaniani, Farnaz, et al.
Published: (2026)
Multi-PA: A Multi-perspective Benchmark on Privacy Assessment for Large Vision-Language Models
by: Zhang, Jie, et al.
Published: (2024)
by: Zhang, Jie, et al.
Published: (2024)
SPICED: Syntactical Bug and Trojan Pattern Identification in A/MS Circuits using LLM-Enhanced Detection
by: Chaudhuri, Jayeeta, et al.
Published: (2024)
by: Chaudhuri, Jayeeta, et al.
Published: (2024)
Benchmarking Large Language Models for Zero-shot and Few-shot Phishing URL Detection
by: Hasan, Najmul, et al.
Published: (2026)
by: Hasan, Najmul, et al.
Published: (2026)
A Study of Vulnerability Repair in JavaScript Programs with Large Language Models
by: Le, Tan Khang, et al.
Published: (2024)
by: Le, Tan Khang, et al.
Published: (2024)
BugSweeper: Function-Level Detection of Smart Contract Vulnerabilities Using Graph Neural Networks
by: Lee, Uisang, et al.
Published: (2025)
by: Lee, Uisang, et al.
Published: (2025)
VulDetectBench: Evaluating the Deep Capability of Vulnerability Detection with Large Language Models
by: Liu, Yu, et al.
Published: (2024)
by: Liu, Yu, et al.
Published: (2024)
VADER: A Human-Evaluated Benchmark for Vulnerability Assessment, Detection, Explanation, and Remediation
by: Liu, Ethan TS., et al.
Published: (2025)
by: Liu, Ethan TS., et al.
Published: (2025)
PrompTrend: Continuous Community-Driven Vulnerability Discovery and Assessment for Large Language Models
by: Gasmi, Tarek, et al.
Published: (2025)
by: Gasmi, Tarek, et al.
Published: (2025)
Can Transformer Memory Be Corrupted? Investigating Cache-Side Vulnerabilities in Large Language Models
by: Hossain, Elias, et al.
Published: (2025)
by: Hossain, Elias, et al.
Published: (2025)
Code Vulnerability Repair with Large Language Model using Context-Aware Prompt Tuning
by: Khan, Arshiya, et al.
Published: (2024)
by: Khan, Arshiya, et al.
Published: (2024)
Jailbreaking and Mitigation of Vulnerabilities in Large Language Models
by: Peng, Benji, et al.
Published: (2024)
by: Peng, Benji, et al.
Published: (2024)
SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code
by: Li, Xinghang, et al.
Published: (2025)
by: Li, Xinghang, et al.
Published: (2025)
Conflicts Make Large Reasoning Models Vulnerable to Attacks
by: Liu, Honghao, et al.
Published: (2026)
by: Liu, Honghao, et al.
Published: (2026)
Behind the Mask: Benchmarking Camouflaged Jailbreaks in Large Language Models
by: Zheng, Youjia, et al.
Published: (2025)
by: Zheng, Youjia, et al.
Published: (2025)
Cyber Threat Detection and Vulnerability Assessment System using Generative AI and Large Language Model
by: M, Keerthi Kumar., et al.
Published: (2026)
by: M, Keerthi Kumar., et al.
Published: (2026)
Pattern Enhanced Multi-Turn Jailbreaking: Exploiting Structural Vulnerabilities in Large Language Models
by: Nihal, Ragib Amin, et al.
Published: (2025)
by: Nihal, Ragib Amin, et al.
Published: (2025)
Distilling Lightweight Language Models for C/C++ Vulnerabilities
by: Wei, Zhiyuan, et al.
Published: (2025)
by: Wei, Zhiyuan, et al.
Published: (2025)
White-Basilisk: A Hybrid Model for Code Vulnerability Detection
by: Lamprou, Ioannis, et al.
Published: (2025)
by: Lamprou, Ioannis, et al.
Published: (2025)
SAEL: Leveraging Large Language Models with Adaptive Mixture-of-Experts for Smart Contract Vulnerability Detection
by: Yu, Lei, et al.
Published: (2025)
by: Yu, Lei, et al.
Published: (2025)
ChatBug: A Common Vulnerability of Aligned LLMs Induced by Chat Templates
by: Jiang, Fengqing, et al.
Published: (2024)
by: Jiang, Fengqing, et al.
Published: (2024)
Analysis of Threat-Based Manipulation in Large Language Models: A Dual Perspective on Vulnerabilities and Performance Enhancement Opportunities
by: Samancioglu, Atil
Published: (2025)
by: Samancioglu, Atil
Published: (2025)
Beyond A Fixed Seal: Adaptive Stealing Watermark in Large Language Models
by: Zhang, Shuhao, et al.
Published: (2026)
by: Zhang, Shuhao, et al.
Published: (2026)
Similar Items
-
BugWhisperer: Fine-Tuning LLMs for SoC Hardware Vulnerability Detection
by: Tarek, Shams, et al.
Published: (2025) -
Beyond Content Safety: Real-Time Monitoring for Reasoning Vulnerabilities in Large Language Models
by: Wang, Xunguang, et al.
Published: (2026) -
Enhancing Reverse Engineering: Investigating and Benchmarking Large Language Models for Vulnerability Analysis in Decompiled Binaries
by: Manuel, Dylan, et al.
Published: (2024) -
Harnessing Large Language Models for Software Vulnerability Detection: A Comprehensive Benchmarking Study
by: Tamberg, Karl, et al.
Published: (2024) -
MARVEL: Multi-Agent RTL Vulnerability Extraction using Large Language Models
by: Collini, Luca, et al.
Published: (2025)