Guardado en:
| Autores principales: | Feiglin, Jake, Dar, Guy |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2601.02941 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
WAInjectBench: Benchmarking Prompt Injection Detections for Web Agents
por: Liu, Yinuo, et al.
Publicado: (2025)
por: Liu, Yinuo, et al.
Publicado: (2025)
RAT-Bench: A Comprehensive Benchmark for Text Anonymization
por: Krčo, Nataša, et al.
Publicado: (2026)
por: Krčo, Nataša, et al.
Publicado: (2026)
A Survey on Agentic Security: Applications, Threats and Defenses
por: Shahriar, Asif, et al.
Publicado: (2025)
por: Shahriar, Asif, et al.
Publicado: (2025)
SALAD-Bench: A Hierarchical and Comprehensive Safety Benchmark for Large Language Models
por: Li, Lijun, et al.
Publicado: (2024)
por: Li, Lijun, et al.
Publicado: (2024)
Text2VLM: Adapting Text-Only Datasets to Evaluate Alignment Training in Visual Language Models
por: Downer, Gabriel, et al.
Publicado: (2025)
por: Downer, Gabriel, et al.
Publicado: (2025)
RedacBench: Can AI Erase Your Secrets?
por: Jeon, Hyunjun, et al.
Publicado: (2026)
por: Jeon, Hyunjun, et al.
Publicado: (2026)
What Was Your Prompt? A Remote Keylogging Attack on AI Assistants
por: Weiss, Roy, et al.
Publicado: (2024)
por: Weiss, Roy, et al.
Publicado: (2024)
AgentSOC: A Multi-Layer Agentic AI Framework for Security Operations Automation
por: Roy, Joyjit, et al.
Publicado: (2026)
por: Roy, Joyjit, et al.
Publicado: (2026)
PoisonBench: Assessing Large Language Model Vulnerability to Data Poisoning
por: Fu, Tingchen, et al.
Publicado: (2024)
por: Fu, Tingchen, et al.
Publicado: (2024)
ExCyTIn-Bench: Evaluating LLM agents on Cyber Threat Investigation
por: Wu, Yiran, et al.
Publicado: (2025)
por: Wu, Yiran, et al.
Publicado: (2025)
Swiss-Bench 003: Evaluating LLM Reliability and Adversarial Security for Swiss Regulatory Contexts
por: Uenal, Fatih
Publicado: (2026)
por: Uenal, Fatih
Publicado: (2026)
Breaking MCP with Function Hijacking Attacks: Novel Threats for Function Calling and Agentic Models
por: Belkhiter, Yannis, et al.
Publicado: (2026)
por: Belkhiter, Yannis, et al.
Publicado: (2026)
AgentAlign: Navigating Safety Alignment in the Shift from Informative to Agentic Large Language Models
por: Zhang, Jinchuan, et al.
Publicado: (2025)
por: Zhang, Jinchuan, et al.
Publicado: (2025)
SecureVibeBench: Benchmarking Secure Vibe Coding of AI Agents via Reconstructing Vulnerability-Introducing Scenarios
por: Chen, Junkai, et al.
Publicado: (2025)
por: Chen, Junkai, et al.
Publicado: (2025)
ChineseHarm-Bench: A Chinese Harmful Content Detection Benchmark
por: Liu, Kangwei, et al.
Publicado: (2025)
por: Liu, Kangwei, et al.
Publicado: (2025)
BELLS: A Framework Towards Future Proof Benchmarks for the Evaluation of LLM Safeguards
por: Dorn, Diego, et al.
Publicado: (2024)
por: Dorn, Diego, et al.
Publicado: (2024)
Bag of Tricks: Benchmarking of Jailbreak Attacks on LLMs
por: Xu, Zhao, et al.
Publicado: (2024)
por: Xu, Zhao, et al.
Publicado: (2024)
Evading Toxicity Detection with ASCII-art: A Benchmark of Spatial Attacks on Moderation Systems
por: Berezin, Sergey, et al.
Publicado: (2024)
por: Berezin, Sergey, et al.
Publicado: (2024)
Agent Tools Orchestration Leaks More: Dataset, Benchmark, and Mitigation
por: Qiao, Yuxuan, et al.
Publicado: (2025)
por: Qiao, Yuxuan, et al.
Publicado: (2025)
SciSafeEval: A Comprehensive Benchmark for Safety Alignment of Large Language Models in Scientific Tasks
por: Li, Tianhao, et al.
Publicado: (2024)
por: Li, Tianhao, et al.
Publicado: (2024)
Test-Time Immunization: A Universal Defense Framework Against Jailbreaks for (Multimodal) Large Language Models
por: Yu, Yongcan, et al.
Publicado: (2025)
por: Yu, Yongcan, et al.
Publicado: (2025)
JailBreakV: A Benchmark for Assessing the Robustness of MultiModal Large Language Models against Jailbreak Attacks
por: Luo, Weidi, et al.
Publicado: (2024)
por: Luo, Weidi, et al.
Publicado: (2024)
Tool Preferences in Agentic LLMs are Unreliable
por: Faghih, Kazem, et al.
Publicado: (2025)
por: Faghih, Kazem, et al.
Publicado: (2025)
Attestable Audits: Verifiable AI Safety Benchmarks Using Trusted Execution Environments
por: Schnabl, Christoph, et al.
Publicado: (2025)
por: Schnabl, Christoph, et al.
Publicado: (2025)
InjecGuard: Benchmarking and Mitigating Over-defense in Prompt Injection Guardrail Models
por: Li, Hao, et al.
Publicado: (2024)
por: Li, Hao, et al.
Publicado: (2024)
Auditing Black-Box LLM APIs with a Rank-Based Uniformity Test
por: Zhu, Xiaoyuan, et al.
Publicado: (2025)
por: Zhu, Xiaoyuan, et al.
Publicado: (2025)
TaeBench: Improving Quality of Toxic Adversarial Examples
por: Zhu, Xuan, et al.
Publicado: (2024)
por: Zhu, Xuan, et al.
Publicado: (2024)
"Moralized" Multi-Step Jailbreak Prompts: Black-Box Testing of Guardrails in Large Language Models for Verbal Attacks
por: Wang, Libo
Publicado: (2024)
por: Wang, Libo
Publicado: (2024)
Can LLMs Keep a Secret? Testing Privacy Implications of Language Models via Contextual Integrity Theory
por: Mireshghallah, Niloofar, et al.
Publicado: (2023)
por: Mireshghallah, Niloofar, et al.
Publicado: (2023)
PII-Bench: Evaluating Query-Aware Privacy Protection Systems
por: Shen, Hao, et al.
Publicado: (2025)
por: Shen, Hao, et al.
Publicado: (2025)
Holistic Automated Red Teaming for Large Language Models through Top-Down Test Case Generation and Multi-turn Interaction
por: Zhang, Jinchuan, et al.
Publicado: (2024)
por: Zhang, Jinchuan, et al.
Publicado: (2024)
AutoPenBench: Benchmarking Generative Agents for Penetration Testing
por: Gioacchini, Luca, et al.
Publicado: (2024)
por: Gioacchini, Luca, et al.
Publicado: (2024)
ContraFix: Agentic Vulnerability Repair via Differential Runtime Evidence and Skill Reuse
por: Liu, Simiao, et al.
Publicado: (2026)
por: Liu, Simiao, et al.
Publicado: (2026)
BountyBench: Dollar Impact of AI Agent Attackers and Defenders on Real-World Cybersecurity Systems
por: Zhang, Andy K., et al.
Publicado: (2025)
por: Zhang, Andy K., et al.
Publicado: (2025)
Adversarial Intent is a Latent Variable: Stateful Trust Inference for Securing Multimodal Agentic RAG
por: Singh, Inderjeet, et al.
Publicado: (2026)
por: Singh, Inderjeet, et al.
Publicado: (2026)
SoK: Agentic Retrieval-Augmented Generation (RAG): Taxonomy, Architectures, Evaluation, and Research Directions
por: Mishra, Saroj, et al.
Publicado: (2026)
por: Mishra, Saroj, et al.
Publicado: (2026)
AgentRedBench: Dynamic Redteaming and Integration-Aware Defense for LLM Agents over SaaS Integrations
por: Dingeto, Hiskias, et al.
Publicado: (2026)
por: Dingeto, Hiskias, et al.
Publicado: (2026)
MCPSecBench: A Systematic Security Benchmark and Playground for Testing Model Context Protocols
por: Yang, Yixuan, et al.
Publicado: (2025)
por: Yang, Yixuan, et al.
Publicado: (2025)
Formalizing and Benchmarking Prompt Injection Attacks and Defenses
por: Liu, Yupei, et al.
Publicado: (2023)
por: Liu, Yupei, et al.
Publicado: (2023)
Effective Red-Teaming of Policy-Adherent Agents
por: Nakash, Itay, et al.
Publicado: (2025)
por: Nakash, Itay, et al.
Publicado: (2025)
Ejemplares similares
-
WAInjectBench: Benchmarking Prompt Injection Detections for Web Agents
por: Liu, Yinuo, et al.
Publicado: (2025) -
RAT-Bench: A Comprehensive Benchmark for Text Anonymization
por: Krčo, Nataša, et al.
Publicado: (2026) -
A Survey on Agentic Security: Applications, Threats and Defenses
por: Shahriar, Asif, et al.
Publicado: (2025) -
SALAD-Bench: A Hierarchical and Comprehensive Safety Benchmark for Large Language Models
por: Li, Lijun, et al.
Publicado: (2024) -
Text2VLM: Adapting Text-Only Datasets to Evaluate Alignment Training in Visual Language Models
por: Downer, Gabriel, et al.
Publicado: (2025)