RAS-Eval: A Comprehensive Benchmark for Security Evaluation of LLM Agents in Real-World Environments
Fuente:
arXiv
Saved in:
| Main Authors: | Fu, Yuchuan, Yuan, Xiaohan, Wang, Dongxia |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Evaluating Privilege Usage of Agents with Real-World Tools
by: Zhang, Quan, et al.
Published: (2026)
by: Zhang, Quan, et al.
Published: (2026)
SecRepoBench: Benchmarking Code Agents for Secure Code Completion in Real-World Repositories
by: Shen, Chihao, et al.
Published: (2025)
by: Shen, Chihao, et al.
Published: (2025)
A New Era in LLM Security: Exploring Security Concerns in Real-World LLM-based Systems
by: Wu, Fangzhou, et al.
Published: (2024)
by: Wu, Fangzhou, et al.
Published: (2024)
Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents
by: Zhang, Hanrong, et al.
Published: (2024)
by: Zhang, Hanrong, et al.
Published: (2024)
Can MLLMs Detect Phishing? A Comprehensive Security Benchmark Suite Focusing on Dynamic Threats and Multimodal Evaluation in Academic Environments
by: Zhou, Jingzhuo
Published: (2025)
by: Zhou, Jingzhuo
Published: (2025)
LogicEval: A Systematic Framework for Evaluating Automated Repair Techniques for Logical Vulnerabilities in Real-World Software
by: Rashid, Syed Md Mukit, et al.
Published: (2026)
by: Rashid, Syed Md Mukit, et al.
Published: (2026)
ENSI: Efficient Non-Interactive Secure Inference for Large Language Models
by: He, Zhiyu, et al.
Published: (2025)
by: He, Zhiyu, et al.
Published: (2025)
A Framework for Formalizing LLM Agent Security
by: Siu, Vincent, et al.
Published: (2026)
by: Siu, Vincent, et al.
Published: (2026)
Towards Secure Retrieval-Augmented Generation: A Comprehensive Review of Threats, Defenses and Benchmarks
by: Mu, Yanming, et al.
Published: (2026)
by: Mu, Yanming, et al.
Published: (2026)
WebSP-Eval: Evaluating Web Agents on Website Security and Privacy Tasks
by: Ramesh, Guruprasad Viswanathan, et al.
Published: (2026)
by: Ramesh, Guruprasad Viswanathan, et al.
Published: (2026)
SkillTester: Benchmarking Utility and Security of Agent Skills
by: Wang, Leye, et al.
Published: (2026)
by: Wang, Leye, et al.
Published: (2026)
From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World
by: Conde, Pedro, et al.
Published: (2026)
by: Conde, Pedro, et al.
Published: (2026)
Agent Security is a Systems Problem
by: Christodorescu, Mihai, et al.
Published: (2026)
by: Christodorescu, Mihai, et al.
Published: (2026)
Security in LLM-as-a-Judge: A Comprehensive SoK
by: Masoud, Aiman Al, et al.
Published: (2026)
by: Masoud, Aiman Al, et al.
Published: (2026)
MCP Security Bench (MSB): Benchmarking Attacks Against Model Context Protocol in LLM Agents
by: Zhang, Dongsen, et al.
Published: (2025)
by: Zhang, Dongsen, et al.
Published: (2025)
Red-Teaming Agent Execution Contexts: Open-World Security Evaluation on OpenClaw
by: Yao, Hongwei, et al.
Published: (2026)
by: Yao, Hongwei, et al.
Published: (2026)
Towards Effective Offensive Security LLM Agents: Hyperparameter Tuning, LLM as a Judge, and a Lightweight CTF Benchmark
by: Shao, Minghao, et al.
Published: (2025)
by: Shao, Minghao, et al.
Published: (2025)
OSS-CRS: Liberating AIxCC Cyber Reasoning Systems for Real-World Open-Source Security
by: Chin, Andrew, et al.
Published: (2026)
by: Chin, Andrew, et al.
Published: (2026)
Agent Audit: A Security Analysis System for LLM Agent Applications
by: Zhang, Haiyue, et al.
Published: (2026)
by: Zhang, Haiyue, et al.
Published: (2026)
The Granularity Mismatch in Agent Security: Argument-Level Provenance Solves Enforcement and Isolates the LLM Reasoning Bottleneck
by: Fan, Linfeng, et al.
Published: (2026)
by: Fan, Linfeng, et al.
Published: (2026)
ClawTrap: A MITM-Based Red-Teaming Framework for Real-World OpenClaw Security Evaluation
by: Zhao, Haochen, et al.
Published: (2026)
by: Zhao, Haochen, et al.
Published: (2026)
SciSafeEval: A Comprehensive Benchmark for Safety Alignment of Large Language Models in Scientific Tasks
by: Li, Tianhao, et al.
Published: (2024)
by: Li, Tianhao, et al.
Published: (2024)
LLM Agents Should Employ Security Principles
by: Zhang, Kaiyuan, et al.
Published: (2025)
by: Zhang, Kaiyuan, et al.
Published: (2025)
When the Manual Lies: A Realistic Benchmark to Evaluate MCP Poisoning Attacks for LLM Agents
by: Liu, Shi, et al.
Published: (2026)
by: Liu, Shi, et al.
Published: (2026)
Mapping LLM Security Landscapes: A Comprehensive Stakeholder Risk Assessment Proposal
by: Pankajakshan, Rahul, et al.
Published: (2024)
by: Pankajakshan, Rahul, et al.
Published: (2024)
Running in CIRCLE? A Simple Benchmark for LLM Code Interpreter Security
by: Chua, Gabriel
Published: (2025)
by: Chua, Gabriel
Published: (2025)
Federated Large Language Models: Feasibility, Robustness, Security and Future Directions
by: Jiang, Wenhao, et al.
Published: (2025)
by: Jiang, Wenhao, et al.
Published: (2025)
A Comparative Evaluation of AI Agent Security Guardrails
by: Li, Qi, et al.
Published: (2026)
by: Li, Qi, et al.
Published: (2026)
FinVault: Benchmarking Financial Agent Safety in Execution-Grounded Environments
by: Yang, Zhi, et al.
Published: (2026)
by: Yang, Zhi, et al.
Published: (2026)
Towards Unifying Quantitative Security Benchmarking for Multi Agent Systems
by: Sharma, Gauri, et al.
Published: (2025)
by: Sharma, Gauri, et al.
Published: (2025)
WASP: Benchmarking Web Agent Security Against Prompt Injection Attacks
by: Evtimov, Ivan, et al.
Published: (2025)
by: Evtimov, Ivan, et al.
Published: (2025)
Measuring Security Without Fooling Ourselves: Why Benchmarking Agents Is Hard
by: Abdelnabi, Sahar, et al.
Published: (2026)
by: Abdelnabi, Sahar, et al.
Published: (2026)
Securing Agentic AI: A Comprehensive Threat Model and Mitigation Framework for Generative AI Agents
by: Narajala, Vineeth Sai, et al.
Published: (2025)
by: Narajala, Vineeth Sai, et al.
Published: (2025)
Enhancing Privacy in Federated Learning: Secure Aggregation for Real-World Healthcare Applications
by: Taiello, Riccardo, et al.
Published: (2024)
by: Taiello, Riccardo, et al.
Published: (2024)
AgentSys: Secure and Dynamic LLM Agents Through Explicit Hierarchical Memory Management
by: Wen, Ruoyao, et al.
Published: (2026)
by: Wen, Ruoyao, et al.
Published: (2026)
PentestEval: Benchmarking LLM-based Penetration Testing with Modular and Stage-Level Design
by: Yang, Ruozhao, et al.
Published: (2025)
by: Yang, Ruozhao, et al.
Published: (2025)
Security Risks in Tool-Enabled AI Agents: A Systematic Analysis of Privileged Execution Environments
by: Goel, Hardik
Published: (2026)
by: Goel, Hardik
Published: (2026)
USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models
by: Zheng, Baolin, et al.
Published: (2025)
by: Zheng, Baolin, et al.
Published: (2025)
ExploitBench: A Capability Ladder Benchmark for LLM Cybersecurity Agents
by: Lee, Seunghyun, et al.
Published: (2026)
by: Lee, Seunghyun, et al.
Published: (2026)
Benchmarking Practices in LLM-driven Offensive Security: Testbeds, Metrics, and Experiment Design
by: Happe, Andreas, et al.
Published: (2025)
by: Happe, Andreas, et al.
Published: (2025)
Similar Items
-
Evaluating Privilege Usage of Agents with Real-World Tools
by: Zhang, Quan, et al.
Published: (2026) -
SecRepoBench: Benchmarking Code Agents for Secure Code Completion in Real-World Repositories
by: Shen, Chihao, et al.
Published: (2025) -
A New Era in LLM Security: Exploring Security Concerns in Real-World LLM-based Systems
by: Wu, Fangzhou, et al.
Published: (2024) -
Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents
by: Zhang, Hanrong, et al.
Published: (2024) -
Can MLLMs Detect Phishing? A Comprehensive Security Benchmark Suite Focusing on Dynamic Threats and Multimodal Evaluation in Academic Environments
by: Zhou, Jingzhuo
Published: (2025)