How to evaluate control measures for LLM agents? A trajectory from today to superintelligence
Fuente:
arXiv
Saved in:
| Main Authors: | Korbak, Tomek, Balesni, Mikita, Shlegeris, Buck, Irving, Geoffrey |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A sketch of an AI control safety case
by: Korbak, Tomek, et al.
Published: (2025)
by: Korbak, Tomek, et al.
Published: (2025)
Towards evaluations-based safety cases for AI scheming
by: Balesni, Mikita, et al.
Published: (2024)
by: Balesni, Mikita, et al.
Published: (2024)
Practical challenges of control monitoring in frontier AI deployments
by: Lindner, David, et al.
Published: (2025)
by: Lindner, David, et al.
Published: (2025)
Lessons from Studying Two-Hop Latent Reasoning
by: Balesni, Mikita, et al.
Published: (2024)
by: Balesni, Mikita, et al.
Published: (2024)
Safety case template for frontier AI: A cyber inability argument
by: Goemans, Arthur, et al.
Published: (2024)
by: Goemans, Arthur, et al.
Published: (2024)
SHADE-Arena: Evaluating Sabotage and Monitoring in LLM Agents
by: Kutasov, Jonathan, et al.
Published: (2025)
by: Kutasov, Jonathan, et al.
Published: (2025)
Security awareness in LLM agents: the NDAI zone case
by: Bottazzi, Enrico, et al.
Published: (2026)
by: Bottazzi, Enrico, et al.
Published: (2026)
AI Kill Switch for malicious web-based LLM agent
by: Lee, Sechan, et al.
Published: (2025)
by: Lee, Sechan, et al.
Published: (2025)
From surveillance to signalling: escalation channels as environmental controls for agentic AI
by: Gomez, Francesca
Published: (2025)
by: Gomez, Francesca
Published: (2025)
How Good LLM-Generated Password Policies Are?
by: Vaidya, Vivek, et al.
Published: (2025)
by: Vaidya, Vivek, et al.
Published: (2025)
Seven Security Challenges That Must be Solved in Cross-domain Multi-agent LLM Systems
by: Ko, Ronny, et al.
Published: (2025)
by: Ko, Ronny, et al.
Published: (2025)
Defense effectiveness across architectural layers: a mechanistic evaluation of persistent memory attacks on stateful LLM agents
by: Leong, Jun Wen
Published: (2026)
by: Leong, Jun Wen
Published: (2026)
Semantic Denial of Service in LLM-controlled robots
by: Steinberg, Jonathan, et al.
Published: (2026)
by: Steinberg, Jonathan, et al.
Published: (2026)
Living Off the LLM: How LLMs Will Change Adversary Tactics
by: Oesch, Sean, et al.
Published: (2025)
by: Oesch, Sean, et al.
Published: (2025)
Next-Generation Phishing: How LLM Agents Empower Cyber Attackers
by: Afane, Khalifa, et al.
Published: (2024)
by: Afane, Khalifa, et al.
Published: (2024)
Robustness of LLM-enabled vehicle trajectory prediction under data security threats
by: Wang, Feilong, et al.
Published: (2025)
by: Wang, Feilong, et al.
Published: (2025)
The System Prompt Is the Attack Surface: How LLM Agent Configuration Shapes Security and Creates Exploitable Vulnerabilities
by: Litvak, Ron
Published: (2026)
by: Litvak, Ron
Published: (2026)
Bergeron: Combating Adversarial Attacks through a Conscience-Based Alignment Framework
by: Pisano, Matthew, et al.
Published: (2023)
by: Pisano, Matthew, et al.
Published: (2023)
How Reliable Are AI Attackers Against a Fixed Vulnerable Target? A 400-Run Empirical Study of LLM Penetration Testing Consistency
by: Erdem, Galip Tolga
Published: (2026)
by: Erdem, Galip Tolga
Published: (2026)
Fundamental Limitations in Pointwise Defences of LLM Finetuning APIs
by: Davies, Xander, et al.
Published: (2025)
by: Davies, Xander, et al.
Published: (2025)
Safety Cases: A Scalable Approach to Frontier AI Safety
by: Hilton, Benjamin, et al.
Published: (2025)
by: Hilton, Benjamin, et al.
Published: (2025)
Context manipulation attacks : Web agents are susceptible to corrupted memory
by: Patlan, Atharv Singh, et al.
Published: (2025)
by: Patlan, Atharv Singh, et al.
Published: (2025)
How Not to Detect Prompt Injections with an LLM
by: Choudhary, Sarthak, et al.
Published: (2025)
by: Choudhary, Sarthak, et al.
Published: (2025)
BadLLM-TG: A Backdoor Defender powered by LLM Trigger Generator
by: Zhang, Ruyi, et al.
Published: (2026)
by: Zhang, Ruyi, et al.
Published: (2026)
Multi-agent Reinforcement Learning-based Network Intrusion Detection System
by: Tellache, Amine, et al.
Published: (2024)
by: Tellache, Amine, et al.
Published: (2024)
ExCyTIn-Bench: Evaluating LLM agents on Cyber Threat Investigation
by: Wu, Yiran, et al.
Published: (2025)
by: Wu, Yiran, et al.
Published: (2025)
LlamaFirewall: An open source guardrail system for building secure AI agents
by: Chennabasappa, Sahana, et al.
Published: (2025)
by: Chennabasappa, Sahana, et al.
Published: (2025)
UNSEEN: A Cross-Stack LLM Unlearning Defense against AR-LLM Social Engineering Attacks
by: Yu, Tianlong, et al.
Published: (2026)
by: Yu, Tianlong, et al.
Published: (2026)
A New Era in LLM Security: Exploring Security Concerns in Real-World LLM-based Systems
by: Wu, Fangzhou, et al.
Published: (2024)
by: Wu, Fangzhou, et al.
Published: (2024)
A Framework for Formalizing LLM Agent Security
by: Siu, Vincent, et al.
Published: (2026)
by: Siu, Vincent, et al.
Published: (2026)
Recursive language models for jailbreak detection: a procedural defense for tool-augmented agents
by: Shavit, Doron
Published: (2026)
by: Shavit, Doron
Published: (2026)
How Catastrophic is Your LLM? Certifying Risk in Conversation
by: Wang, Chengxiao, et al.
Published: (2025)
by: Wang, Chengxiao, et al.
Published: (2025)
Information Security Based on LLM Approaches: A Review
by: Gong, Chang, et al.
Published: (2025)
by: Gong, Chang, et al.
Published: (2025)
SAGE: A Generic Framework for LLM Safety Evaluation
by: Jindal, Madhur, et al.
Published: (2025)
by: Jindal, Madhur, et al.
Published: (2025)
Autonomous LLM Agents & CTFs: A Second Look
by: Bouchari, Youness, et al.
Published: (2026)
by: Bouchari, Youness, et al.
Published: (2026)
The Great Pretender: A Stochasticity Problem in LLM Jailbreak
by: Monteuuis, Jean-Philippe, et al.
Published: (2026)
by: Monteuuis, Jean-Philippe, et al.
Published: (2026)
CheatAgent: Attacking LLM-Empowered Recommender Systems via LLM Agent
by: Ning, Liang-bo, et al.
Published: (2025)
by: Ning, Liang-bo, et al.
Published: (2025)
LLM Access Shield: Domain-Specific LLM Framework for Privacy Policy Compliance
by: Wang, Yu, et al.
Published: (2025)
by: Wang, Yu, et al.
Published: (2025)
VulnLLM-R: Specialized Reasoning LLM with Agent Scaffold for Vulnerability Detection
by: Nie, Yuzhou, et al.
Published: (2025)
by: Nie, Yuzhou, et al.
Published: (2025)
How Far Are VLMs from Privacy Awareness in the Physical World? An Empirical Study
by: Wang, Junran, et al.
Published: (2026)
by: Wang, Junran, et al.
Published: (2026)
Similar Items
-
A sketch of an AI control safety case
by: Korbak, Tomek, et al.
Published: (2025) -
Towards evaluations-based safety cases for AI scheming
by: Balesni, Mikita, et al.
Published: (2024) -
Practical challenges of control monitoring in frontier AI deployments
by: Lindner, David, et al.
Published: (2025) -
Lessons from Studying Two-Hop Latent Reasoning
by: Balesni, Mikita, et al.
Published: (2024) -
Safety case template for frontier AI: A cyber inability argument
by: Goemans, Arthur, et al.
Published: (2024)