Out of Distribution, Out of Luck: How Well Can LLMs Trained on Vulnerability Datasets Detect Top 25 CWE Weaknesses?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Yikun, Bui, Ngoc Tan, Zhang, Ting, Yang, Chengran, Zhou, Xin, Weyssow, Martin, Jiang, Jinfeng, Chen, Junkai, Huang, Huihui, Nguyen, Huu Hung, Ho, Chiok Yew, Tan, Jie, Li, Ruiyin, Yin, Yide, Ang, Han Wei, Liauw, Frank, Ouh, Eng Lieh, Shar, Lwin Khin, Lo, David |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
R2Vul: Learning to Reason about Software Vulnerabilities with Reinforcement Learning and Structured Reasoning Distillation
von: Weyssow, Martin, et al.
Veröffentlicht: (2025)
von: Weyssow, Martin, et al.
Veröffentlicht: (2025)
PenForge: On-the-Fly Expert Agent Construction for Automated Penetration Testing
von: Huang, Huihui, et al.
Veröffentlicht: (2026)
von: Huang, Huihui, et al.
Veröffentlicht: (2026)
Benchmarking Large Language Models for Multi-Language Software Vulnerability Detection
von: Zhang, Ting, et al.
Veröffentlicht: (2025)
von: Zhang, Ting, et al.
Veröffentlicht: (2025)
An Execution-Verified Multi-Language Benchmark for Code Semantic Reasoning
von: Li, Yikun, et al.
Veröffentlicht: (2026)
von: Li, Yikun, et al.
Veröffentlicht: (2026)
Let the Trial Begin: A Mock-Court Approach to Vulnerability Detection using LLM-Based Agents
von: Widyasari, Ratnadira, et al.
Veröffentlicht: (2025)
von: Widyasari, Ratnadira, et al.
Veröffentlicht: (2025)
PatchSeeker: Mapping NVD Records to their Vulnerability-fixing Commits with LLM Generated Commits and Embeddings
von: Nguyen, Huu Hung, et al.
Veröffentlicht: (2025)
von: Nguyen, Huu Hung, et al.
Veröffentlicht: (2025)
VulCoCo: A Simple Yet Effective Method for Detecting Vulnerable Code Clones
von: Bui, Tan, et al.
Veröffentlicht: (2025)
von: Bui, Tan, et al.
Veröffentlicht: (2025)
Beyond Function-Level Analysis: Context-Aware Reasoning for Inter-Procedural Vulnerability Detection
von: Li, Yikun, et al.
Veröffentlicht: (2026)
von: Li, Yikun, et al.
Veröffentlicht: (2026)
Semantics-Aligned, Curriculum-Driven, and Reasoning-Enhanced Vulnerability Repair Framework
von: Yang, Chengran, et al.
Veröffentlicht: (2025)
von: Yang, Chengran, et al.
Veröffentlicht: (2025)
CleanVul: Automatic Function-Level Vulnerability Detection in Code Commits Using LLM Heuristics
von: Li, Yikun, et al.
Veröffentlicht: (2024)
von: Li, Yikun, et al.
Veröffentlicht: (2024)
Back to the Basics: Rethinking Issue-Commit Linking with LLM-Assisted Retrieval
von: Huang, Huihui, et al.
Veröffentlicht: (2025)
von: Huang, Huihui, et al.
Veröffentlicht: (2025)
TitanCA: Lessons from Orchestrating LLM Agents to Discover 100+ CVEs
von: Zhang, Ting, et al.
Veröffentlicht: (2026)
von: Zhang, Ting, et al.
Veröffentlicht: (2026)
Revisiting Vulnerability Patch Identification on Data in the Wild
von: Irsan, Ivana Clairine, et al.
Veröffentlicht: (2026)
von: Irsan, Ivana Clairine, et al.
Veröffentlicht: (2026)
Mapping NVD Records to Their Vulnerability-fixing Commits: How Hard is It?
von: Nguyen, Huu Hung, et al.
Veröffentlicht: (2025)
von: Nguyen, Huu Hung, et al.
Veröffentlicht: (2025)
Security Modelling for Cyber-Physical Systems: A Systematic Literature Review
von: Huang, Shaofei, et al.
Veröffentlicht: (2024)
von: Huang, Shaofei, et al.
Veröffentlicht: (2024)
Bayesian and Multi-Objective Decision Support for Real-Time Incident Mitigation in Critical Infrastructure
von: Huang, Shaofei, et al.
Veröffentlicht: (2025)
von: Huang, Shaofei, et al.
Veröffentlicht: (2025)
From Incomplete Architecture to Quantified Risk: Multimodal LLM-Driven Security Assessment for Cyber-Physical Systems
von: Huang, Shaofei, et al.
Veröffentlicht: (2026)
von: Huang, Shaofei, et al.
Veröffentlicht: (2026)
ACTISM: Threat-informed Dynamic Security Modelling for Automotive Systems
von: Huang, Shaofei, et al.
Veröffentlicht: (2024)
von: Huang, Shaofei, et al.
Veröffentlicht: (2024)
Runtime Anomaly Detection for Drones: An Integrated Rule-Mining and Unsupervised-Learning Approach
von: Tan, Ivan, et al.
Veröffentlicht: (2025)
von: Tan, Ivan, et al.
Veröffentlicht: (2025)
The Price of Prompting: Profiling Energy Use in Large Language Models Inference
von: Husom, Erik Johannes, et al.
Veröffentlicht: (2024)
von: Husom, Erik Johannes, et al.
Veröffentlicht: (2024)
VLM-Fuzz: Vision Language Model Assisted Recursive Depth-first Search Exploration for Effective UI Testing of Android Apps
von: Demissie, Biniam Fisseha, et al.
Veröffentlicht: (2025)
von: Demissie, Biniam Fisseha, et al.
Veröffentlicht: (2025)
Beginner's Luck Has Just Run Out.
von: Loch-Wouters, Marge
Veröffentlicht: (1991)
von: Loch-Wouters, Marge
Veröffentlicht: (1991)
Luck Out or Outpay? Competing with a Public Option
von: Mekonnen, Teddy
Veröffentlicht: (2025)
von: Mekonnen, Teddy
Veröffentlicht: (2025)
Differentiated Security Architecture for Secure and Efficient Infotainment Data Communication in IoV Networks
von: Fan, Jiani, et al.
Veröffentlicht: (2024)
von: Fan, Jiani, et al.
Veröffentlicht: (2024)
On the Usage of Continual Learning for Out-of-Distribution Generalization in Pre-trained Language Models of Code
von: Weyssow, Martin, et al.
Veröffentlicht: (2023)
von: Weyssow, Martin, et al.
Veröffentlicht: (2023)
Decentralized Multimedia Data Sharing in IoV: A Learning-based Equilibrium of Supply and Demand
von: Fan, Jiani, et al.
Veröffentlicht: (2024)
von: Fan, Jiani, et al.
Veröffentlicht: (2024)
Towards Reliable LLM-Driven Fuzz Testing: Vision and Road Ahead
von: Cheng, Yiran, et al.
Veröffentlicht: (2025)
von: Cheng, Yiran, et al.
Veröffentlicht: (2025)
A Model-Constrained Discontinuous Galerkin Network (DGNet) for Compressible Euler Equations with Out-of-Distribution Generalization
von: Nguyen, Hai V., et al.
Veröffentlicht: (2024)
von: Nguyen, Hai V., et al.
Veröffentlicht: (2024)
CovAgent: Overcoming the 30% Curse of Mobile Application Coverage with Agentic AI and Dynamic Instrumentation
von: Minn, Wei, et al.
Veröffentlicht: (2026)
von: Minn, Wei, et al.
Veröffentlicht: (2026)
ESTANDARIZACIÓN Y VALIDACIÓN DE LA TÉCNICA RT-PCR CUALITATIVA EN TIEMPO REAL PARA LA DETECCIÓN DEL VIRUS DE LA PESTE PORCINA CLÁSICA
von: Kim Lam Chiok C.
Veröffentlicht: (2011)
von: Kim Lam Chiok C.
Veröffentlicht: (2011)
LIDL: LLM Integration Defect Localization via Knowledge Graph-Enhanced Multi-Agent Analysis
von: Tan, Gou, et al.
Veröffentlicht: (2026)
von: Tan, Gou, et al.
Veröffentlicht: (2026)
Deep Learning Approaches for Anti-Money Laundering on Mobile Transactions: Review, Framework, and Directions
von: Fan, Jiani, et al.
Veröffentlicht: (2025)
von: Fan, Jiani, et al.
Veröffentlicht: (2025)
Sustainable LLM Inference for Edge AI: Evaluating Quantized LLMs for Energy Efficiency, Output Accuracy, and Inference Latency
von: Husom, Erik Johannes, et al.
Veröffentlicht: (2025)
von: Husom, Erik Johannes, et al.
Veröffentlicht: (2025)
Fixseeker: An Empirical Driven Graph-based Approach for Detecting Silent Vulnerability Fixes in Open Source Software
von: Cheng, Yiran, et al.
Veröffentlicht: (2025)
von: Cheng, Yiran, et al.
Veröffentlicht: (2025)
VERCATION: Precise Vulnerable Open-source Software Version Identification based on Static Analysis and LLM
von: Cheng, Yiran, et al.
Veröffentlicht: (2024)
von: Cheng, Yiran, et al.
Veröffentlicht: (2024)
Resilience Analysis of k‐Out‐of‐n Systems
von: Bei Wu, et al.
Veröffentlicht: (2025)
von: Bei Wu, et al.
Veröffentlicht: (2025)
Beyond the Tip of the Iceberg: Understanding SATD in Dockerfiles through the Lens of Co-evolution
von: Minn, Wei, et al.
Veröffentlicht: (2026)
von: Minn, Wei, et al.
Veröffentlicht: (2026)
Bamboo: LLM-Driven Discovery of API-Permission Mappings in the Android Framework
von: Hu, Han, et al.
Veröffentlicht: (2025)
von: Hu, Han, et al.
Veröffentlicht: (2025)
“Good Luck Out There Without NDIS”: Challenges Accessing Individualized Support Packages by Autistic Young People Leaving School
von: Caroline Mills, et al.
Veröffentlicht: (2025)
von: Caroline Mills, et al.
Veröffentlicht: (2025)
Feature Protection For Out-of-distribution Generalization
von: Tan, Lu, et al.
Veröffentlicht: (2024)
von: Tan, Lu, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
R2Vul: Learning to Reason about Software Vulnerabilities with Reinforcement Learning and Structured Reasoning Distillation
von: Weyssow, Martin, et al.
Veröffentlicht: (2025) -
PenForge: On-the-Fly Expert Agent Construction for Automated Penetration Testing
von: Huang, Huihui, et al.
Veröffentlicht: (2026) -
Benchmarking Large Language Models for Multi-Language Software Vulnerability Detection
von: Zhang, Ting, et al.
Veröffentlicht: (2025) -
An Execution-Verified Multi-Language Benchmark for Code Semantic Reasoning
von: Li, Yikun, et al.
Veröffentlicht: (2026) -
Let the Trial Begin: A Mock-Court Approach to Vulnerability Detection using LLM-Based Agents
von: Widyasari, Ratnadira, et al.
Veröffentlicht: (2025)