SEC-bench: Automated Benchmarking of LLM Agents on Real-World Software Security Tasks
Fuente:
arXiv
Saved in:
| Main Authors: | Lee, Hwiwon, Zhang, Ziqi, Lu, Hanxiao, Zhang, Lingming |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SEC-bench Pro: Can Language Models Solve Long-Horizon Software Security Tasks?
by: Lee, Hwiwon, et al.
Published: (2026)
by: Lee, Hwiwon, et al.
Published: (2026)
Agentic Vulnerability Reasoning on Windows COM Binaries
by: Lee, Hwiwon, et al.
Published: (2026)
by: Lee, Hwiwon, et al.
Published: (2026)
How Safe Are AI-Generated Patches? A Large-scale Study on Security Risks in LLM and Agentic Automated Program Repair on SWE-bench
by: Sajadi, Amirali, et al.
Published: (2025)
by: Sajadi, Amirali, et al.
Published: (2025)
CTFusion: A CTF-based Benchmark for LLM Agent Evaluation
by: Lee, Dongjun, et al.
Published: (2026)
by: Lee, Dongjun, et al.
Published: (2026)
Automated Detection and Analysis of Data Practices Using A Real-World Corpus
by: Srinath, Mukund, et al.
Published: (2024)
by: Srinath, Mukund, et al.
Published: (2024)
RealSec-bench: A Benchmark for Evaluating Secure Code Generation in Real-World Repositories
by: Wang, Yanlin, et al.
Published: (2026)
by: Wang, Yanlin, et al.
Published: (2026)
Co-RedTeam: Orchestrated Security Discovery and Exploitation with LLM Agents
by: He, Pengfei, et al.
Published: (2026)
by: He, Pengfei, et al.
Published: (2026)
Design Patterns for Securing LLM Agents against Prompt Injections
by: Beurer-Kellner, Luca, et al.
Published: (2025)
by: Beurer-Kellner, Luca, et al.
Published: (2025)
Benchmarking LLAMA Model Security Against OWASP Top 10 For LLM Applications
by: Shahin, Nourin, et al.
Published: (2026)
by: Shahin, Nourin, et al.
Published: (2026)
MCPTox: A Benchmark for Tool Poisoning Attack on Real-World MCP Servers
by: Wang, Zhiqiang, et al.
Published: (2025)
by: Wang, Zhiqiang, et al.
Published: (2025)
From Trace to Line: LLM Agent for Real-World OSS Vulnerability Localization
by: Xi, Haoran, et al.
Published: (2025)
by: Xi, Haoran, et al.
Published: (2025)
Breaking Agents: Compromising Autonomous LLM Agents Through Malfunction Amplification
by: Zhang, Boyang, et al.
Published: (2024)
by: Zhang, Boyang, et al.
Published: (2024)
Towards a Generalisable Cyber Defence Agent for Real-World Computer Networks
by: Dudman, Tim, et al.
Published: (2025)
by: Dudman, Tim, et al.
Published: (2025)
Automated Membership Inference Attacks: Discovering MIA Signal Computations using LLM Agents
by: Tran, Toan, et al.
Published: (2026)
by: Tran, Toan, et al.
Published: (2026)
SecureInfer: Heterogeneous TEE-GPU Architecture for Privacy-Critical Tensors for Large Language Model Deployment
by: Nayan, Tushar, et al.
Published: (2025)
by: Nayan, Tushar, et al.
Published: (2025)
A Safety and Security Framework for Real-World Agentic Systems
by: Ghosh, Shaona, et al.
Published: (2025)
by: Ghosh, Shaona, et al.
Published: (2025)
CyberGym: Evaluating AI Agents' Real-World Cybersecurity Capabilities at Scale
by: Wang, Zhun, et al.
Published: (2025)
by: Wang, Zhun, et al.
Published: (2025)
What Do They Fix? LLM-Aided Categorization of Security Patches for Critical Memory Bugs
by: Li, Xingyu, et al.
Published: (2025)
by: Li, Xingyu, et al.
Published: (2025)
A Systematic Survey of Security Threats and Defenses in LLM-Based AI Agents: A Layered Attack Surface Framework
by: Chu, Kexin
Published: (2026)
by: Chu, Kexin
Published: (2026)
Testbed and Software Architecture for Enhancing Security in Industrial Private 5G Networks
by: Ha, Song Son, et al.
Published: (2025)
by: Ha, Song Son, et al.
Published: (2025)
AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents
by: Debenedetti, Edoardo, et al.
Published: (2024)
by: Debenedetti, Edoardo, et al.
Published: (2024)
HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns
by: Shen, Xinyue, et al.
Published: (2025)
by: Shen, Xinyue, et al.
Published: (2025)
Information Security and Privacy in the Digital World: Some Selected Topics
by: Sen, Jaydip, et al.
Published: (2024)
by: Sen, Jaydip, et al.
Published: (2024)
Security Considerations for Artificial Intelligence Agents
by: Li, Ninghui, et al.
Published: (2026)
by: Li, Ninghui, et al.
Published: (2026)
WebSP-Eval: Evaluating Web Agents on Website Security and Privacy Tasks
by: Ramesh, Guruprasad Viswanathan, et al.
Published: (2026)
by: Ramesh, Guruprasad Viswanathan, et al.
Published: (2026)
MPC-Minimized Secure LLM Inference
by: Rathee, Deevashwer, et al.
Published: (2024)
by: Rathee, Deevashwer, et al.
Published: (2024)
FraudFox: Adaptable Fraud Detection in the Real World
by: Butler, Matthew, et al.
Published: (2026)
by: Butler, Matthew, et al.
Published: (2026)
ACE: A Security Architecture for LLM-Integrated App Systems
by: Li, Evan, et al.
Published: (2025)
by: Li, Evan, et al.
Published: (2025)
Assessing Deanonymization Risks with Stylometry-Assisted LLM Agent
by: Zhang, Boyang, et al.
Published: (2026)
by: Zhang, Boyang, et al.
Published: (2026)
ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks?
by: Wang, Zhun, et al.
Published: (2026)
by: Wang, Zhun, et al.
Published: (2026)
Secure Transformer Inference Protocol
by: Yuan, Mu, et al.
Published: (2023)
by: Yuan, Mu, et al.
Published: (2023)
Federated Learning for Cross-Domain Data Privacy: A Distributed Approach to Secure Collaboration
by: Zhang, Yiwei, et al.
Published: (2025)
by: Zhang, Yiwei, et al.
Published: (2025)
Real-World Adversarial Attacks on RF-Based Drone Detectors
by: Gazit, Omer, et al.
Published: (2025)
by: Gazit, Omer, et al.
Published: (2025)
LLM Security and Safety: Insights from Homotopy-Inspired Prompt Obfuscation
by: Lazo, Luis, et al.
Published: (2026)
by: Lazo, Luis, et al.
Published: (2026)
A Unified Compliance Aggregator Framework for Automated Multi-Tool Security Assessment of Linux Systems
by: Paul, Sheldon, et al.
Published: (2026)
by: Paul, Sheldon, et al.
Published: (2026)
Automated Cyber Defense with Generalizable Graph-based Reinforcement Learning Agents
by: King, Isaiah J., et al.
Published: (2025)
by: King, Isaiah J., et al.
Published: (2025)
Memory-Induced Tool-Drift in LLM Agents
by: Dabas, Mahavir, et al.
Published: (2026)
by: Dabas, Mahavir, et al.
Published: (2026)
Shell or Nothing: Real-World Benchmarks and Memory-Activated Agents for Automated Penetration Testing
by: Mai, Wuyuao, et al.
Published: (2025)
by: Mai, Wuyuao, et al.
Published: (2025)
Private Federated Learning In Real World Application -- A Case Study
by: Ji, An, et al.
Published: (2025)
by: Ji, An, et al.
Published: (2025)
MAIDS: Malicious Agent Identification-based Data Security Model for Cloud Environments
by: Gupta, Kishu, et al.
Published: (2024)
by: Gupta, Kishu, et al.
Published: (2024)
Similar Items
-
SEC-bench Pro: Can Language Models Solve Long-Horizon Software Security Tasks?
by: Lee, Hwiwon, et al.
Published: (2026) -
Agentic Vulnerability Reasoning on Windows COM Binaries
by: Lee, Hwiwon, et al.
Published: (2026) -
How Safe Are AI-Generated Patches? A Large-scale Study on Security Risks in LLM and Agentic Automated Program Repair on SWE-bench
by: Sajadi, Amirali, et al.
Published: (2025) -
CTFusion: A CTF-based Benchmark for LLM Agent Evaluation
by: Lee, Dongjun, et al.
Published: (2026) -
Automated Detection and Analysis of Data Practices Using A Real-World Corpus
by: Srinath, Mukund, et al.
Published: (2024)