Gespeichert in:
| Hauptverfasser: | Jotautaitė, Monika, Martinez, Maria Angelica, Matthews, Ollie, Tracy, Tyler |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2605.09684 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SLEIGHT-Bench: A Benchmark of Evasion Attacks Against Agent Monitors
von: Najt, Elle, et al.
Veröffentlicht: (2026)
von: Najt, Elle, et al.
Veröffentlicht: (2026)
RepliBench: Evaluating the Autonomous Replication Capabilities of Language Model Agents
von: Black, Sid, et al.
Veröffentlicht: (2025)
von: Black, Sid, et al.
Veröffentlicht: (2025)
SHADE-Arena: Evaluating Sabotage and Monitoring in LLM Agents
von: Kutasov, Jonathan, et al.
Veröffentlicht: (2025)
von: Kutasov, Jonathan, et al.
Veröffentlicht: (2025)
SafeSearch: Automated Red-Teaming of LLM-Based Search Agents
von: Dong, Jianshuo, et al.
Veröffentlicht: (2025)
von: Dong, Jianshuo, et al.
Veröffentlicht: (2025)
Attack Selection Reduces Safety in Concentrated AI Control Settings against Trusted Monitoring
von: Schaeffer, Joachim, et al.
Veröffentlicht: (2026)
von: Schaeffer, Joachim, et al.
Veröffentlicht: (2026)
Red Teaming AI Red Teaming
von: Majumdar, Subhabrata, et al.
Veröffentlicht: (2025)
von: Majumdar, Subhabrata, et al.
Veröffentlicht: (2025)
Proteus: A Self-Evolving Red Team for Agent Skill Ecosystems
von: Zhou, Zhaojiacheng
Veröffentlicht: (2026)
von: Zhou, Zhaojiacheng
Veröffentlicht: (2026)
BashArena: A Control Setting for Highly Privileged AI Agents
von: Kaufman, Adam, et al.
Veröffentlicht: (2025)
von: Kaufman, Adam, et al.
Veröffentlicht: (2025)
LLM Agent Honeypot: Monitoring AI Hacking Agents in the Wild
von: Reworr, et al.
Veröffentlicht: (2024)
von: Reworr, et al.
Veröffentlicht: (2024)
Red Teaming Large Reasoning Models
von: Chen, Jiawei, et al.
Veröffentlicht: (2025)
von: Chen, Jiawei, et al.
Veröffentlicht: (2025)
Whispers of Wealth: Red-Teaming Google's Agent Payments Protocol via Prompt Injection
von: Debi, Tanusree, et al.
Veröffentlicht: (2026)
von: Debi, Tanusree, et al.
Veröffentlicht: (2026)
Red-Teaming Agent Execution Contexts: Open-World Security Evaluation on OpenClaw
von: Yao, Hongwei, et al.
Veröffentlicht: (2026)
von: Yao, Hongwei, et al.
Veröffentlicht: (2026)
RedAgent: Red Teaming Large Language Models with Context-aware Autonomous Language Agent
von: Xu, Huiyu, et al.
Veröffentlicht: (2024)
von: Xu, Huiyu, et al.
Veröffentlicht: (2024)
AutoRedTeamer: Autonomous Red Teaming with Lifelong Attack Integration
von: Zhou, Andy, et al.
Veröffentlicht: (2025)
von: Zhou, Andy, et al.
Veröffentlicht: (2025)
MUZZLE: Adaptive Agentic Red-Teaming of Web Agents Against Indirect Prompt Injection Attacks
von: Syros, Georgios, et al.
Veröffentlicht: (2026)
von: Syros, Georgios, et al.
Veröffentlicht: (2026)
Red-Teaming Coding Agents from a Tool-Invocation Perspective: An Empirical Security Assessment
von: Xie, Yuchong, et al.
Veröffentlicht: (2025)
von: Xie, Yuchong, et al.
Veröffentlicht: (2025)
Is Monitoring Enough? Strategic Agent Selection For Stealthy Attack in Multi-Agent Discussions
von: Xiang, Qiuchi, et al.
Veröffentlicht: (2026)
von: Xiang, Qiuchi, et al.
Veröffentlicht: (2026)
Monitoring Decomposition Attacks in LLMs with Lightweight Sequential Monitors
von: Yueh-Han, Chen, et al.
Veröffentlicht: (2025)
von: Yueh-Han, Chen, et al.
Veröffentlicht: (2025)
Adaptive Instruction Composition for Automated LLM Red-Teaming
von: Zymet, Jesse, et al.
Veröffentlicht: (2026)
von: Zymet, Jesse, et al.
Veröffentlicht: (2026)
IPI-proxy: An Intercepting Proxy for Red-Teaming Web-Browsing AI Agents Against Indirect Prompt Injection
von: Chia-Pei, et al.
Veröffentlicht: (2026)
von: Chia-Pei, et al.
Veröffentlicht: (2026)
T-MAP: Red-Teaming LLM Agents with Trajectory-aware Evolutionary Search
von: Lee, Hyomin, et al.
Veröffentlicht: (2026)
von: Lee, Hyomin, et al.
Veröffentlicht: (2026)
Effective Red-Teaming of Policy-Adherent Agents
von: Nakash, Itay, et al.
Veröffentlicht: (2025)
von: Nakash, Itay, et al.
Veröffentlicht: (2025)
A Red Teaming Roadmap Towards System-Level Safety
von: Wang, Zifan, et al.
Veröffentlicht: (2025)
von: Wang, Zifan, et al.
Veröffentlicht: (2025)
Automatic Red Teaming LLM-based Agents with Model Context Protocol Tools
von: He, Ping, et al.
Veröffentlicht: (2025)
von: He, Ping, et al.
Veröffentlicht: (2025)
Reliable Weak-to-Strong Monitoring of LLM Agents
von: Kale, Neil, et al.
Veröffentlicht: (2025)
von: Kale, Neil, et al.
Veröffentlicht: (2025)
Redefining AI Red Teaming in the Agentic Era: From Weeks to Hours
von: Dheekonda, Raja Sekhar Rao, et al.
Veröffentlicht: (2026)
von: Dheekonda, Raja Sekhar Rao, et al.
Veröffentlicht: (2026)
An Application-Layer Multi-Modal Covert-Channel Reference Monitor for LLM Agent Egress
von: Metere, Alfredo
Veröffentlicht: (2026)
von: Metere, Alfredo
Veröffentlicht: (2026)
From Firewalls to Frontiers: AI Red-Teaming is a Domain-Specific Evolution of Cyber Red-Teaming
von: Sinha, Anusha, et al.
Veröffentlicht: (2025)
von: Sinha, Anusha, et al.
Veröffentlicht: (2025)
Blue Teaming Function-Calling Agents
von: Dolcetti, Greta, et al.
Veröffentlicht: (2026)
von: Dolcetti, Greta, et al.
Veröffentlicht: (2026)
AutoRISE: Agent-Driven Strategy Evolution for Red-Teaming Large Language Models
von: Gautam, Tanmay, et al.
Veröffentlicht: (2026)
von: Gautam, Tanmay, et al.
Veröffentlicht: (2026)
SIRAJ: Diverse and Efficient Red-Teaming for LLM Agents via Distilled Structured Reasoning
von: Zhou, Kaiwen, et al.
Veröffentlicht: (2025)
von: Zhou, Kaiwen, et al.
Veröffentlicht: (2025)
Incalmo: An Autonomous LLM-assisted System for Red Teaming Multi-Host Networks
von: Singer, Brian, et al.
Veröffentlicht: (2025)
von: Singer, Brian, et al.
Veröffentlicht: (2025)
BlackIce: A Containerized Red Teaming Toolkit for AI Security Testing
von: Kaplan, Caelin, et al.
Veröffentlicht: (2025)
von: Kaplan, Caelin, et al.
Veröffentlicht: (2025)
A Systematic Review of Algorithmic Red Teaming Methodologies for Assurance and Security of AI Applications
von: Srivastava, Shruti, et al.
Veröffentlicht: (2026)
von: Srivastava, Shruti, et al.
Veröffentlicht: (2026)
When Search Goes Wrong: Red-Teaming Web-Augmented Large Language Models
von: Ou, Haoran, et al.
Veröffentlicht: (2025)
von: Ou, Haoran, et al.
Veröffentlicht: (2025)
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training
von: Du, Pengfei
Veröffentlicht: (2025)
von: Du, Pengfei
Veröffentlicht: (2025)
LumiMAS: A Comprehensive Framework for Real-Time Monitoring and Enhanced Observability in Multi-Agent Systems
von: Solomon, Ron, et al.
Veröffentlicht: (2025)
von: Solomon, Ron, et al.
Veröffentlicht: (2025)
CoT-Guard: Small Models for Strong Monitoring
von: Diwan, Nirav, et al.
Veröffentlicht: (2026)
von: Diwan, Nirav, et al.
Veröffentlicht: (2026)
Red-Teaming Claude Opus and ChatGPT-based Security Advisors for Trusted Execution Environments
von: Mukherjee, Kunal, et al.
Veröffentlicht: (2026)
von: Mukherjee, Kunal, et al.
Veröffentlicht: (2026)
CORVUS: Red-Teaming Hallucination Detectors via Internal Signal Camouflage in Large Language Models
von: Min, Nay Myat, et al.
Veröffentlicht: (2026)
von: Min, Nay Myat, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
SLEIGHT-Bench: A Benchmark of Evasion Attacks Against Agent Monitors
von: Najt, Elle, et al.
Veröffentlicht: (2026) -
RepliBench: Evaluating the Autonomous Replication Capabilities of Language Model Agents
von: Black, Sid, et al.
Veröffentlicht: (2025) -
SHADE-Arena: Evaluating Sabotage and Monitoring in LLM Agents
von: Kutasov, Jonathan, et al.
Veröffentlicht: (2025) -
SafeSearch: Automated Red-Teaming of LLM-Based Search Agents
von: Dong, Jianshuo, et al.
Veröffentlicht: (2025) -
Attack Selection Reduces Safety in Concentrated AI Control Settings against Trusted Monitoring
von: Schaeffer, Joachim, et al.
Veröffentlicht: (2026)