SEC-bench Pro: Can Language Models Solve Long-Horizon Software Security Tasks?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lee, Hwiwon, Liu, Jiawei, Kim, Dongjun, Zhang, Ziqi, Xia, Chunqiu Steven, Zhang, Lingming |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SEC-bench: Automated Benchmarking of LLM Agents on Real-World Software Security Tasks
von: Lee, Hwiwon, et al.
Veröffentlicht: (2025)
von: Lee, Hwiwon, et al.
Veröffentlicht: (2025)
Agentic Vulnerability Reasoning on Windows COM Binaries
von: Lee, Hwiwon, et al.
Veröffentlicht: (2026)
von: Lee, Hwiwon, et al.
Veröffentlicht: (2026)
SLIP-SEC: Formalizing Secure Protocols for Model IP Protection
von: Jain, Racchit, et al.
Veröffentlicht: (2025)
von: Jain, Racchit, et al.
Veröffentlicht: (2025)
KernelGPT: Enhanced Kernel Fuzzing via Large Language Models
von: Yang, Chenyuan, et al.
Veröffentlicht: (2023)
von: Yang, Chenyuan, et al.
Veröffentlicht: (2023)
Establishing a Baseline of Software Supply Chain Security Task Adoption by Software Organizations
von: Williams, Laurie, et al.
Veröffentlicht: (2025)
von: Williams, Laurie, et al.
Veröffentlicht: (2025)
RealSec-bench: A Benchmark for Evaluating Secure Code Generation in Real-World Repositories
von: Wang, Yanlin, et al.
Veröffentlicht: (2026)
von: Wang, Yanlin, et al.
Veröffentlicht: (2026)
Exploring Advanced Methodologies in Security Evaluation for LLMs
von: Huang, Jun, et al.
Veröffentlicht: (2024)
von: Huang, Jun, et al.
Veröffentlicht: (2024)
Can Small Language Models Reliably Resist Jailbreak Attacks? A Comprehensive Evaluation
von: Zhang, Wenhui, et al.
Veröffentlicht: (2025)
von: Zhang, Wenhui, et al.
Veröffentlicht: (2025)
Hybrid Horizons: Policy for Post-Quantum Security
von: Jaikissoon, Anais
Veröffentlicht: (2025)
von: Jaikissoon, Anais
Veröffentlicht: (2025)
SecureInfer: Heterogeneous TEE-GPU Architecture for Privacy-Critical Tensors for Large Language Model Deployment
von: Nayan, Tushar, et al.
Veröffentlicht: (2025)
von: Nayan, Tushar, et al.
Veröffentlicht: (2025)
Risk Assessment and Security Analysis of Large Language Models
von: Zhang, Xiaoyan, et al.
Veröffentlicht: (2025)
von: Zhang, Xiaoyan, et al.
Veröffentlicht: (2025)
SQL Injection Jailbreak: A Structural Disaster of Large Language Models
von: Zhao, Jiawei, et al.
Veröffentlicht: (2024)
von: Zhao, Jiawei, et al.
Veröffentlicht: (2024)
An Empirical Evaluation of LLMs for Solving Offensive Security Challenges
von: Shao, Minghao, et al.
Veröffentlicht: (2024)
von: Shao, Minghao, et al.
Veröffentlicht: (2024)
Time-Frequency Jointed Imperceptible Adversarial Attack to Brainprint Recognition with Deep Learning Models
von: Yi, Hangjie, et al.
Veröffentlicht: (2024)
von: Yi, Hangjie, et al.
Veröffentlicht: (2024)
AnyPoC: Universal Proof-of-Concept Test Generation for Scalable LLM-Based Bug Detection
von: Zhao, Zijie, et al.
Veröffentlicht: (2026)
von: Zhao, Zijie, et al.
Veröffentlicht: (2026)
Bridging Theory and Practice: An Executable Taxonomy of Security Properties for ProVerif and Tamarin
von: Tudorache, Leonard, et al.
Veröffentlicht: (2026)
von: Tudorache, Leonard, et al.
Veröffentlicht: (2026)
Weaver: Fuzzing JavaScript Engines at the JavaScript-WebAssembly Boundary
von: Zhang, Lingming, et al.
Veröffentlicht: (2026)
von: Zhang, Lingming, et al.
Veröffentlicht: (2026)
Enhancing Software Supply Chain Resilience: Strategy For Mitigating Software Supply Chain Security Risks And Ensuring Security Continuity In Development Lifecycle
von: Akinsola, Ahmed, et al.
Veröffentlicht: (2024)
von: Akinsola, Ahmed, et al.
Veröffentlicht: (2024)
STEAD: Robust Provably Secure Linguistic Steganography with Diffusion Language Model
von: Qi, Yuang, et al.
Veröffentlicht: (2026)
von: Qi, Yuang, et al.
Veröffentlicht: (2026)
Software-Defined Cryptography: A Design Feature of Cryptographic Agility
von: Cho, Jihoon, et al.
Veröffentlicht: (2024)
von: Cho, Jihoon, et al.
Veröffentlicht: (2024)
Software Unclonable Functions for IoT Devices Identification and Security
von: Alshehhi, Saeed
Veröffentlicht: (2025)
von: Alshehhi, Saeed
Veröffentlicht: (2025)
An Empirical Study on Virtual Reality Software Security Weaknesses
von: Xu, Yifan, et al.
Veröffentlicht: (2025)
von: Xu, Yifan, et al.
Veröffentlicht: (2025)
Security Vulnerabilities in Software Supply Chain for Autonomous Vehicles
von: Haque, Md Wasiul, et al.
Veröffentlicht: (2025)
von: Haque, Md Wasiul, et al.
Veröffentlicht: (2025)
Large Language Models for Blockchain Security: A Systematic Literature Review
von: He, Zheyuan, et al.
Veröffentlicht: (2024)
von: He, Zheyuan, et al.
Veröffentlicht: (2024)
Large Language Models for Cyber Security
von: Somani, Raunak, et al.
Veröffentlicht: (2025)
von: Somani, Raunak, et al.
Veröffentlicht: (2025)
Assessing the Software Security Comprehension of Large Language Models
von: Siddiq, Mohammed Latif, et al.
Veröffentlicht: (2025)
von: Siddiq, Mohammed Latif, et al.
Veröffentlicht: (2025)
How Can We Effectively Use LLMs for Phishing Detection?: Evaluating the Effectiveness of Large Language Model-based Phishing Detection Models
von: Ji, Fujiao, et al.
Veröffentlicht: (2025)
von: Ji, Fujiao, et al.
Veröffentlicht: (2025)
Practical Secure Inference Algorithm for Fine-tuned Large Language Model Based on Fully Homomorphic Encryption
von: Ruoyan, Zhang, et al.
Veröffentlicht: (2025)
von: Ruoyan, Zhang, et al.
Veröffentlicht: (2025)
Adjacent Words, Divergent Intents: Jailbreaking Large Language Models via Task Concurrency
von: Jiang, Yukun, et al.
Veröffentlicht: (2025)
von: Jiang, Yukun, et al.
Veröffentlicht: (2025)
ALPS: Automated Least-Privilege Enforcement for Securing Serverless Functions
von: Shin, Changhee, et al.
Veröffentlicht: (2026)
von: Shin, Changhee, et al.
Veröffentlicht: (2026)
ProSec: Fortifying Code LLMs with Proactive Security Alignment
von: Xu, Xiangzhe, et al.
Veröffentlicht: (2024)
von: Xu, Xiangzhe, et al.
Veröffentlicht: (2024)
The Hidden Dangers of Outdated Software: A Cyber Security Perspective
von: Thiyagarajan, Gogulakrishnan, et al.
Veröffentlicht: (2025)
von: Thiyagarajan, Gogulakrishnan, et al.
Veröffentlicht: (2025)
Formal Security Analysis of the AMD SEV-SNP Software Interface
von: Paradžik, Petar, et al.
Veröffentlicht: (2024)
von: Paradžik, Petar, et al.
Veröffentlicht: (2024)
An Approach for Safe and Secure Software Protection Supported by Symbolic Execution
von: Dorfmeister, Daniel, et al.
Veröffentlicht: (2026)
von: Dorfmeister, Daniel, et al.
Veröffentlicht: (2026)
CTFusion: A CTF-based Benchmark for LLM Agent Evaluation
von: Lee, Dongjun, et al.
Veröffentlicht: (2026)
von: Lee, Dongjun, et al.
Veröffentlicht: (2026)
Silent Guardian: Protecting Text from Malicious Exploitation by Large Language Models
von: Zhao, Jiawei, et al.
Veröffentlicht: (2023)
von: Zhao, Jiawei, et al.
Veröffentlicht: (2023)
CAPTCHA Solving for Native GUI Agents: Automated Reasoning-Action Data Generation and Self-Corrective Training
von: Chen, Yuxi, et al.
Veröffentlicht: (2026)
von: Chen, Yuxi, et al.
Veröffentlicht: (2026)
CS-Eval: A Comprehensive Large Language Model Benchmark for CyberSecurity
von: Yu, Zhengmin, et al.
Veröffentlicht: (2024)
von: Yu, Zhengmin, et al.
Veröffentlicht: (2024)
Large Language Model (LLM) for Software Security: Code Analysis, Malware Analysis, Reverse Engineering
von: Jelodar, Hamed, et al.
Veröffentlicht: (2025)
von: Jelodar, Hamed, et al.
Veröffentlicht: (2025)
Security and Privacy on Generative Data in AIGC: A Survey
von: Wang, Tao, et al.
Veröffentlicht: (2023)
von: Wang, Tao, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
SEC-bench: Automated Benchmarking of LLM Agents on Real-World Software Security Tasks
von: Lee, Hwiwon, et al.
Veröffentlicht: (2025) -
Agentic Vulnerability Reasoning on Windows COM Binaries
von: Lee, Hwiwon, et al.
Veröffentlicht: (2026) -
SLIP-SEC: Formalizing Secure Protocols for Model IP Protection
von: Jain, Racchit, et al.
Veröffentlicht: (2025) -
KernelGPT: Enhanced Kernel Fuzzing via Large Language Models
von: Yang, Chenyuan, et al.
Veröffentlicht: (2023) -
Establishing a Baseline of Software Supply Chain Security Task Adoption by Software Organizations
von: Williams, Laurie, et al.
Veröffentlicht: (2025)