Evaluating Control Protocols for Untrusted AI Agents
Fuente:
arXiv
Guardado en:
| Autores principales: | Kutasov, Jon, Loughridge, Chloe, Sun, Yuqi, Sleight, Henry, Shlegeris, Buck, Tracy, Tyler, Benton, Joe |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Optimizing AI Agent Attacks With Synthetic Data
por: Loughridge, Chloe, et al.
Publicado: (2025)
por: Loughridge, Chloe, et al.
Publicado: (2025)
SHADE-Arena: Evaluating Sabotage and Monitoring in LLM Agents
por: Kutasov, Jonathan, et al.
Publicado: (2025)
por: Kutasov, Jonathan, et al.
Publicado: (2025)
Games for AI Control: Models of Safety Evaluations of AI Deployment Protocols
por: Griffin, Charlie, et al.
Publicado: (2024)
por: Griffin, Charlie, et al.
Publicado: (2024)
Polysemanticity and Capacity in Neural Networks
por: Scherlis, Adam, et al.
Publicado: (2022)
por: Scherlis, Adam, et al.
Publicado: (2022)
SLEIGHT-Bench: A Benchmark of Evasion Attacks Against Agent Monitors
por: Najt, Elle, et al.
Publicado: (2026)
por: Najt, Elle, et al.
Publicado: (2026)
dafny-annotator: AI-Assisted Verification of Dafny Programs
por: Poesia, Gabriel, et al.
Publicado: (2024)
por: Poesia, Gabriel, et al.
Publicado: (2024)
Model Spec Midtraining: Improving How Alignment Training Generalizes
por: Li, Chloe, et al.
Publicado: (2026)
por: Li, Chloe, et al.
Publicado: (2026)
A sketch of an AI control safety case
por: Korbak, Tomek, et al.
Publicado: (2025)
por: Korbak, Tomek, et al.
Publicado: (2025)
How to evaluate control measures for LLM agents? A trajectory from today to superintelligence
por: Korbak, Tomek, et al.
Publicado: (2025)
por: Korbak, Tomek, et al.
Publicado: (2025)
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases
por: Gan, Eric, et al.
Publicado: (2026)
por: Gan, Eric, et al.
Publicado: (2026)
Factor(U,T): Controlling Untrusted AI by Monitoring their Plans
por: Lip, Edward Lue Chee, et al.
Publicado: (2025)
por: Lip, Edward Lue Chee, et al.
Publicado: (2025)
Safe, Untrusted, "Proof-Carrying" AI Agents: toward the agentic lakehouse
por: Tagliabue, Jacopo, et al.
Publicado: (2025)
por: Tagliabue, Jacopo, et al.
Publicado: (2025)
BashArena: A Control Setting for Highly Privileged AI Agents
por: Kaufman, Adam, et al.
Publicado: (2025)
por: Kaufman, Adam, et al.
Publicado: (2025)
Subversion Strategy Eval: Can language models statelessly strategize to subvert control protocols?
por: Mallen, Alex, et al.
Publicado: (2024)
por: Mallen, Alex, et al.
Publicado: (2024)
Language models are better than humans at next-token prediction
por: Shlegeris, Buck, et al.
Publicado: (2022)
por: Shlegeris, Buck, et al.
Publicado: (2022)
Adaptive Deployment of Untrusted LLMs Reduces Distributed Threats
por: Wen, Jiaxin, et al.
Publicado: (2024)
por: Wen, Jiaxin, et al.
Publicado: (2024)
Efficiently Aligning Language Models with Online Natural Language Feedback
por: Ye, Christine, et al.
Publicado: (2026)
por: Ye, Christine, et al.
Publicado: (2026)
Sabotage Evaluations for Frontier Models
por: Benton, Joe, et al.
Publicado: (2024)
por: Benton, Joe, et al.
Publicado: (2024)
Ctrl-Z: Controlling AI Agents via Resampling
por: Bhatt, Aryan, et al.
Publicado: (2025)
por: Bhatt, Aryan, et al.
Publicado: (2025)
AI Organizations are More Effective but Less Aligned than Individual Agents
por: Shen, Judy Hanwen, et al.
Publicado: (2026)
por: Shen, Judy Hanwen, et al.
Publicado: (2026)
The Hot Mess of AI: How Does Misalignment Scale With Model Intelligence and Task Complexity?
por: Hägele, Alexander, et al.
Publicado: (2026)
por: Hägele, Alexander, et al.
Publicado: (2026)
Singular Vectors of Attention Heads Align with Features
por: Franco, Gabriel, et al.
Publicado: (2026)
por: Franco, Gabriel, et al.
Publicado: (2026)
Structured Security Auditing and Robustness Enhancement for Untrusted Agent Skills
por: Lv, Lijia, et al.
Publicado: (2026)
por: Lv, Lijia, et al.
Publicado: (2026)
All Code, No Thought: Current Language Models Struggle to Reason in Ciphered Language
por: Guo, Shiyuan, et al.
Publicado: (2025)
por: Guo, Shiyuan, et al.
Publicado: (2025)
The LLM Has Left The Chat: Evidence of Bail Preferences in Large Language Models
por: Ensign, Danielle, et al.
Publicado: (2025)
por: Ensign, Danielle, et al.
Publicado: (2025)
Attack Selection Reduces Safety in Concentrated AI Control Settings against Trusted Monitoring
por: Schaeffer, Joachim, et al.
Publicado: (2026)
por: Schaeffer, Joachim, et al.
Publicado: (2026)
A Survey of AI Agent Protocols
por: Yang, Yingxuan, et al.
Publicado: (2025)
por: Yang, Yingxuan, et al.
Publicado: (2025)
LinuxArena: A Control Setting for AI Agents in Live Production Software Environments
por: Tracy, Tyler, et al.
Publicado: (2026)
por: Tracy, Tyler, et al.
Publicado: (2026)
Bypassing AI Control Protocols via Agent-as-a-Proxy Attacks
por: Isbarov, Jafar, et al.
Publicado: (2026)
por: Isbarov, Jafar, et al.
Publicado: (2026)
Agent Control Protocol: Admission Control for Agent Actions
por: Fernandez, Marcelo
Publicado: (2026)
por: Fernandez, Marcelo
Publicado: (2026)
MonitoringBench: Semi-Automated Red-Teaming for Agent Monitoring
por: Jotautaitė, Monika, et al.
Publicado: (2026)
por: Jotautaitė, Monika, et al.
Publicado: (2026)
PragLocker: Protecting Agent Intellectual Property in Untrusted Deployments via Non-Portable Prompts
por: Li, Qinfeng, et al.
Publicado: (2026)
por: Li, Qinfeng, et al.
Publicado: (2026)
Stress-Testing Model Specs Reveals Character Differences among Language Models
por: Zhang, Jifan, et al.
Publicado: (2025)
por: Zhang, Jifan, et al.
Publicado: (2025)
Data-Centric Interpretability for LLM-based Multi-Agent Reinforcement Learning
por: Yan, John, et al.
Publicado: (2026)
por: Yan, John, et al.
Publicado: (2026)
Deep Research Bench: Evaluating AI Web Research Agents
por: FutureSearch, et al.
Publicado: (2025)
por: FutureSearch, et al.
Publicado: (2025)
From Prompts to Protocols: An AI Agent for Laboratory Automation
por: Angelopoulos, Angelos, et al.
Publicado: (2026)
por: Angelopoulos, Angelos, et al.
Publicado: (2026)
Inverse Scaling in Test-Time Compute
por: Gema, Aryo Pradipta, et al.
Publicado: (2025)
por: Gema, Aryo Pradipta, et al.
Publicado: (2025)
Generative AI in clinical practice: novel qualitative evidence of risk and responsible use of Google's NotebookLM
por: Reuter, Max, et al.
Publicado: (2025)
por: Reuter, Max, et al.
Publicado: (2025)
KGCE: Knowledge-Augmented Dual-Graph Evaluator for Cross-Platform Educational Agent Benchmarking with Multimodal Language Models
por: Liu, Zixian, et al.
Publicado: (2026)
por: Liu, Zixian, et al.
Publicado: (2026)
A survey of agent interoperability protocols: Model Context Protocol (MCP), Agent Communication Protocol (ACP), Agent-to-Agent Protocol (A2A), and Agent Network Protocol (ANP)
por: Ehtesham, Abul, et al.
Publicado: (2025)
por: Ehtesham, Abul, et al.
Publicado: (2025)
Ejemplares similares
-
Optimizing AI Agent Attacks With Synthetic Data
por: Loughridge, Chloe, et al.
Publicado: (2025) -
SHADE-Arena: Evaluating Sabotage and Monitoring in LLM Agents
por: Kutasov, Jonathan, et al.
Publicado: (2025) -
Games for AI Control: Models of Safety Evaluations of AI Deployment Protocols
por: Griffin, Charlie, et al.
Publicado: (2024) -
Polysemanticity and Capacity in Neural Networks
por: Scherlis, Adam, et al.
Publicado: (2022) -
SLEIGHT-Bench: A Benchmark of Evasion Attacks Against Agent Monitors
por: Najt, Elle, et al.
Publicado: (2026)