CASTLE: Benchmarking Dataset for Static Code Analyzers and LLMs towards CWE Detection
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Dubniczky, Richard A., Horvát, Krisztofer Zoltán, Bisztray, Tamás, Ferrag, Mohamed Amine, Cordeiro, Lucas C., Tihanyi, Norbert |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
The Hidden DNA of LLM-Generated JavaScript: Structural Patterns Enable High-Accuracy Authorship Attribution
par: Tihanyi, Norbert, et autres
Publié: (2025)
par: Tihanyi, Norbert, et autres
Publié: (2025)
CyberMetric: A Benchmark Dataset based on Retrieval-Augmented Generation for Evaluating LLMs in Cybersecurity Knowledge
par: Tihanyi, Norbert, et autres
Publié: (2024)
par: Tihanyi, Norbert, et autres
Publié: (2024)
How secure is AI-generated Code: A Large-Scale Comparison of Large Language Models
par: Tihanyi, Norbert, et autres
Publié: (2024)
par: Tihanyi, Norbert, et autres
Publié: (2024)
Vulnerability Detection: From Formal Verification to Large Language Models and Hybrid Approaches: A Comprehensive Overview
par: Tihanyi, Norbert, et autres
Publié: (2025)
par: Tihanyi, Norbert, et autres
Publié: (2025)
Securing Tomorrow's Smart Cities: Investigating Software Security in Internet of Vehicles and Deep Learning Technologies
par: Jain, Ridhi, et autres
Publié: (2024)
par: Jain, Ridhi, et autres
Publié: (2024)
You Have Been LaTeXpOsEd: A Systematic Analysis of Information Leakage in Preprint Archives Using Large Language Models
par: Dubniczky, Richard A., et autres
Publié: (2025)
par: Dubniczky, Richard A., et autres
Publié: (2025)
I Know Which LLM Wrote Your Code Last Summer: LLM generated Code Stylometry for Authorship Attribution
par: Bisztray, Tamas, et autres
Publié: (2025)
par: Bisztray, Tamas, et autres
Publié: (2025)
DFIR-Metric: A Benchmark Dataset for Evaluating Large Language Models in Digital Forensics and Incident Response
par: Cherif, Bilel, et autres
Publié: (2025)
par: Cherif, Bilel, et autres
Publié: (2025)
The FormAI Dataset: Generative AI in Software Security Through the Lens of Formal Verification
par: Tihanyi, Norbert, et autres
Publié: (2023)
par: Tihanyi, Norbert, et autres
Publié: (2023)
Generative AI in Cybersecurity: A Comprehensive Review of LLM Applications and Vulnerabilities
par: Ferrag, Mohamed Amine, et autres
Publié: (2024)
par: Ferrag, Mohamed Amine, et autres
Publié: (2024)
A New Era in Software Security: Towards Self-Healing Software via Large Language Models and Formal Verification
par: Tihanyi, Norbert, et autres
Publié: (2023)
par: Tihanyi, Norbert, et autres
Publié: (2023)
Out of Distribution, Out of Luck: How Well Can LLMs Trained on Vulnerability Datasets Detect Top 25 CWE Weaknesses?
par: Li, Yikun, et autres
Publié: (2025)
par: Li, Yikun, et autres
Publié: (2025)
Revolutionizing Cyber Threat Detection with Large Language Models: A privacy-preserving BERT-based Lightweight Model for IoT/IIoT Devices
par: Ferrag, Mohamed Amine, et autres
Publié: (2023)
par: Ferrag, Mohamed Amine, et autres
Publié: (2023)
Edge Learning for 6G-enabled Internet of Things: A Comprehensive Survey of Vulnerabilities, Datasets, and Defenses
par: Ferrag, Mohamed Amine, et autres
Publié: (2023)
par: Ferrag, Mohamed Amine, et autres
Publié: (2023)
Sustaining Cyber Awareness: The Long-Term Impact of Continuous Phishing Training and Emotional Triggers
par: Toth, Rebeka, et autres
Publié: (2025)
par: Toth, Rebeka, et autres
Publié: (2025)
From Generalist to Specialist: Exploring CWE-Specific Vulnerability Detection
par: Atiiq, Syafiq Al, et autres
Publié: (2024)
par: Atiiq, Syafiq Al, et autres
Publié: (2024)
From Prompt Injections to Protocol Exploits: Threats in LLM-Powered AI Agents Workflows
par: Ferrag, Mohamed Amine, et autres
Publié: (2025)
par: Ferrag, Mohamed Amine, et autres
Publié: (2025)
Many Tools, Few Exploitable Vulnerabilities: A Survey of 246 Static Code Analyzers for Security
par: Hermann, Kevin, et autres
Publié: (2026)
par: Hermann, Kevin, et autres
Publié: (2026)
Think Broad, Act Narrow: CWE Identification with Multi-Agent Large Language Models
par: Sayagh, Mohammed, et autres
Publié: (2025)
par: Sayagh, Mohammed, et autres
Publié: (2025)
LLMs in Web Development: Evaluating LLM-Generated PHP Code Unveiling Vulnerabilities and Limitations
par: Tóth, Rebeka, et autres
Publié: (2024)
par: Tóth, Rebeka, et autres
Publié: (2024)
Reasoning Beyond Limits: Advances and Open Problems for LLMs
par: Ferrag, Mohamed Amine, et autres
Publié: (2025)
par: Ferrag, Mohamed Amine, et autres
Publié: (2025)
Dynamic Intelligence Assessment: Benchmarking LLMs on the Road to AGI with a Focus on Model Confidence
par: Tihanyi, Norbert, et autres
Publié: (2024)
par: Tihanyi, Norbert, et autres
Publié: (2024)
The Phish, The Spam, and The Valid: Generating Feature-Rich Emails for Benchmarking LLMs
par: Toth, Rebeka, et autres
Publié: (2025)
par: Toth, Rebeka, et autres
Publié: (2025)
SecureFalcon: Are We There Yet in Automated Software Vulnerability Detection with LLMs?
par: Ferrag, Mohamed Amine, et autres
Publié: (2023)
par: Ferrag, Mohamed Amine, et autres
Publié: (2023)
DITING: A Static Analyzer for Identifying Bad Partitioning Issues in TEE Applications
par: Ma, Chengyan, et autres
Publié: (2025)
par: Ma, Chengyan, et autres
Publié: (2025)
Guiding Symbolic Execution with Static Analysis and LLMs for Vulnerability Discovery
par: Shafiuzzaman, Md, et autres
Publié: (2026)
par: Shafiuzzaman, Md, et autres
Publié: (2026)
From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review
par: Ferrag, Mohamed Amine, et autres
Publié: (2025)
par: Ferrag, Mohamed Amine, et autres
Publié: (2025)
SCRIBE: Practical Static Binary Patching via Binary-Aware Recompilation of Decompiled Code
par: Dai, Han, et autres
Publié: (2026)
par: Dai, Han, et autres
Publié: (2026)
Towards Secure Logging: Characterizing and Benchmarking Logging Code Security Issues with LLMs
par: Yuan, He Yang, et autres
Publié: (2026)
par: Yuan, He Yang, et autres
Publié: (2026)
QLPro: Automated Code Vulnerability Discovery via LLM and Static Code Analysis Integration
par: Hu, Junze, et autres
Publié: (2025)
par: Hu, Junze, et autres
Publié: (2025)
Can I Check What I Designed? Mapping Security Design DSLs to Code Analyzers
par: Peldszus, Sven, et autres
Publié: (2026)
par: Peldszus, Sven, et autres
Publié: (2026)
CrossCommitVuln-Bench: A Dataset of Multi-Commit Python Vulnerabilities Invisible to Per-Commit Static Analysis
par: Majumdar, Arunabh
Publié: (2026)
par: Majumdar, Arunabh
Publié: (2026)
DeCoMa: Detecting and Purifying Code Dataset Watermarks through Dual Channel Code Abstraction
par: Xiao, Yuan, et autres
Publié: (2025)
par: Xiao, Yuan, et autres
Publié: (2025)
An Empirical Security Evaluation of LLM-Generated Cryptographic Rust Code
par: Elsayed, Mohamed, et autres
Publié: (2026)
par: Elsayed, Mohamed, et autres
Publié: (2026)
Decoding Secret Memorization in Code LLMs Through Token-Level Characterization
par: Nie, Yuqing, et autres
Publié: (2024)
par: Nie, Yuqing, et autres
Publié: (2024)
Does Teaming-Up LLMs Improve Secure Code Generation? A Comprehensive Evaluation with Multi-LLMSecCodeEval
par: Sabir, Bushra, et autres
Publié: (2026)
par: Sabir, Bushra, et autres
Publié: (2026)
Train in Vain: Functionality-Preserving Poisoning to Prevent Unauthorized Use of Code Datasets
par: Xiao, Yuan, et autres
Publié: (2026)
par: Xiao, Yuan, et autres
Publié: (2026)
MegaVul: A C/C++ Vulnerability Dataset with Comprehensive Code Representation
par: Ni, Chao, et autres
Publié: (2024)
par: Ni, Chao, et autres
Publié: (2024)
CodableLLM: Automating Decompiled and Source Code Mapping for LLM Dataset Generation
par: Manuel, Dylan, et autres
Publié: (2025)
par: Manuel, Dylan, et autres
Publié: (2025)
Hybrid Privacy Policy-Code Consistency Check using Knowledge Graphs and LLMs
par: Mao, Zhenyu, et autres
Publié: (2025)
par: Mao, Zhenyu, et autres
Publié: (2025)
Documents similaires
-
The Hidden DNA of LLM-Generated JavaScript: Structural Patterns Enable High-Accuracy Authorship Attribution
par: Tihanyi, Norbert, et autres
Publié: (2025) -
CyberMetric: A Benchmark Dataset based on Retrieval-Augmented Generation for Evaluating LLMs in Cybersecurity Knowledge
par: Tihanyi, Norbert, et autres
Publié: (2024) -
How secure is AI-generated Code: A Large-Scale Comparison of Large Language Models
par: Tihanyi, Norbert, et autres
Publié: (2024) -
Vulnerability Detection: From Formal Verification to Large Language Models and Hybrid Approaches: A Comprehensive Overview
par: Tihanyi, Norbert, et autres
Publié: (2025) -
Securing Tomorrow's Smart Cities: Investigating Software Security in Internet of Vehicles and Deep Learning Technologies
par: Jain, Ridhi, et autres
Publié: (2024)