A sketch of an AI control safety case
Fuente:
arXiv
Saved in:
| Main Authors: | Korbak, Tomek, Clymer, Joshua, Hilton, Benjamin, Shlegeris, Buck, Irving, Geoffrey |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
How to evaluate control measures for LLM agents? A trajectory from today to superintelligence
by: Korbak, Tomek, et al.
Published: (2025)
by: Korbak, Tomek, et al.
Published: (2025)
Towards evaluations-based safety cases for AI scheming
by: Balesni, Mikita, et al.
Published: (2024)
by: Balesni, Mikita, et al.
Published: (2024)
LinuxArena: A Control Setting for AI Agents in Live Production Software Environments
by: Tracy, Tyler, et al.
Published: (2026)
by: Tracy, Tyler, et al.
Published: (2026)
Practical challenges of control monitoring in frontier AI deployments
by: Lindner, David, et al.
Published: (2025)
by: Lindner, David, et al.
Published: (2025)
LLMs + Security = Trouble
by: Livshits, Benjamin
Published: (2026)
by: Livshits, Benjamin
Published: (2026)
Deconstructing Obfuscation: A four-dimensional framework for evaluating Large Language Models assembly code deobfuscation capabilities
by: Tkachenko, Anton, et al.
Published: (2025)
by: Tkachenko, Anton, et al.
Published: (2025)
Securing the Future of IVR: AI-Driven Innovation with Agile Security, Data Regulation, and Ethical AI Integration
by: Shaikh, Khushbu Mehboob, et al.
Published: (2025)
by: Shaikh, Khushbu Mehboob, et al.
Published: (2025)
Broken by Default: A Formal Verification Study of Security Vulnerabilities in AI-Generated Code
by: Blain, Dominik, et al.
Published: (2026)
by: Blain, Dominik, et al.
Published: (2026)
AI security and cyber risk in IoT systems
by: Radanliev, Petar, et al.
Published: (2024)
by: Radanliev, Petar, et al.
Published: (2024)
AI Code Generators for Security: Friend or Foe?
by: Natella, Roberto, et al.
Published: (2024)
by: Natella, Roberto, et al.
Published: (2024)
Containment Verification: AI Safety Guarantees Independent of Alignment
by: Moon, Royce, et al.
Published: (2026)
by: Moon, Royce, et al.
Published: (2026)
Risks of ignoring uncertainty propagation in AI-augmented security pipelines
by: Mezzi, Emanuele, et al.
Published: (2024)
by: Mezzi, Emanuele, et al.
Published: (2024)
Identifying the Supply Chain of AI for Trustworthiness and Risk Management in Critical Applications
by: Sheh, Raymond K., et al.
Published: (2025)
by: Sheh, Raymond K., et al.
Published: (2025)
Data and Context Matter: Towards Generalizing AI-based Software Vulnerability Detection
by: Safdar, Rijha, et al.
Published: (2025)
by: Safdar, Rijha, et al.
Published: (2025)
DevOps-Gym: Benchmarking AI Agents in Software DevOps Cycle
by: Tang, Yuheng, et al.
Published: (2026)
by: Tang, Yuheng, et al.
Published: (2026)
Poisoning Programs by Un-Repairing Code: Security Concerns of AI-generated Code
by: Improta, Cristina
Published: (2024)
by: Improta, Cristina
Published: (2024)
Testing Storage-System Correctness: Challenges, Fuzzing Limitations, and AI-Augmented Opportunities
by: Wang, Ying, et al.
Published: (2026)
by: Wang, Ying, et al.
Published: (2026)
Block MedCare: Advancing healthcare through blockchain integration with AI and IoT
by: Simonoski, Oliver, et al.
Published: (2024)
by: Simonoski, Oliver, et al.
Published: (2024)
Safety case template for frontier AI: A cyber inability argument
by: Goemans, Arthur, et al.
Published: (2024)
by: Goemans, Arthur, et al.
Published: (2024)
AIBoMGen: Generating an AI Bill of Materials for Secure, Transparent, and Compliant Model Training
by: Vandendriessche, Wiebe, et al.
Published: (2026)
by: Vandendriessche, Wiebe, et al.
Published: (2026)
Zer0n: An AI-Assisted Vulnerability Discovery and Blockchain-Backed Integrity Framework
by: Parmar, Harshil, et al.
Published: (2026)
by: Parmar, Harshil, et al.
Published: (2026)
SOK: Exploring Hallucinations and Security Risks in AI-Assisted Software Development with Insights for LLM Deployment
by: Haque, Ariful, et al.
Published: (2025)
by: Haque, Ariful, et al.
Published: (2025)
MalCodeAI: Autonomous Vulnerability Detection and Remediation via Language Agnostic Code Reasoning
by: Gajjar, Jugal, et al.
Published: (2025)
by: Gajjar, Jugal, et al.
Published: (2025)
Investigating Detection and Obfuscation of Prompt Injection Attacks Against Software Reverse Engineering AI Agents
by: Crawford, Brian, et al.
Published: (2026)
by: Crawford, Brian, et al.
Published: (2026)
Safety and Performance, Why Not Both? Bi-Objective Optimized Model Compression against Heterogeneous Attacks Toward AI Software Deployment
by: Zhu, Jie, et al.
Published: (2024)
by: Zhu, Jie, et al.
Published: (2024)
Beyond Embeddings: Interpretable Feature Extraction for Binary Code Similarity
by: Gagnon, Charles E., et al.
Published: (2025)
by: Gagnon, Charles E., et al.
Published: (2025)
Rethinking Autonomy: Preventing Failures in AI-Driven Software Engineering
by: Navneet, Satyam Kumar, et al.
Published: (2025)
by: Navneet, Satyam Kumar, et al.
Published: (2025)
Innamark: A Whitespace Replacement Information-Hiding Method
by: Hellmeier, Malte, et al.
Published: (2025)
by: Hellmeier, Malte, et al.
Published: (2025)
Security of LLM-generated Code: A Comparative Analysis
by: Morkonda, Srivathsan G, et al.
Published: (2026)
by: Morkonda, Srivathsan G, et al.
Published: (2026)
Cochise: A Reference Harness for Autonomous Penetration Testing
by: Happe, Andreas, et al.
Published: (2026)
by: Happe, Andreas, et al.
Published: (2026)
SCDBench: A Benchmark for LLM-Based Smart Contract Decompilers
by: Qin, Kaihua, et al.
Published: (2026)
by: Qin, Kaihua, et al.
Published: (2026)
Graph Neural Networks for Vulnerability Detection: A Counterfactual Explanation
by: Chu, Zhaoyang, et al.
Published: (2024)
by: Chu, Zhaoyang, et al.
Published: (2024)
On the Security Risks of ML-based Malware Detection Systems: A Survey
by: He, Ping, et al.
Published: (2025)
by: He, Ping, et al.
Published: (2025)
A Comprehensive Study of Exploitable Patterns in Smart Contracts: From Vulnerability to Defense
by: Ding, Yuchen, et al.
Published: (2025)
by: Ding, Yuchen, et al.
Published: (2025)
Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing
by: Peng, Jiaren, et al.
Published: (2026)
by: Peng, Jiaren, et al.
Published: (2026)
Understanding Privacy Risks in Code Models Through Training Dynamics: A Causal Approach
by: Yang, Hua, et al.
Published: (2025)
by: Yang, Hua, et al.
Published: (2025)
Harnessing Large Language Models for Software Vulnerability Detection: A Comprehensive Benchmarking Study
by: Tamberg, Karl, et al.
Published: (2024)
by: Tamberg, Karl, et al.
Published: (2024)
A Qualitative Study on Using ChatGPT for Software Security: Perception vs. Practicality
by: Kholoosi, M. Mehdi, et al.
Published: (2024)
by: Kholoosi, M. Mehdi, et al.
Published: (2024)
Measuring the Permission Gate: A Stress-Test Evaluation of Claude Code's Auto Mode
by: Ji, Zimo, et al.
Published: (2026)
by: Ji, Zimo, et al.
Published: (2026)
SecureFixAgent: A Hybrid LLM Agent for Automated Python Static Vulnerability Repair
by: Gajjar, Jugal, et al.
Published: (2025)
by: Gajjar, Jugal, et al.
Published: (2025)
Similar Items
-
How to evaluate control measures for LLM agents? A trajectory from today to superintelligence
by: Korbak, Tomek, et al.
Published: (2025) -
Towards evaluations-based safety cases for AI scheming
by: Balesni, Mikita, et al.
Published: (2024) -
LinuxArena: A Control Setting for AI Agents in Live Production Software Environments
by: Tracy, Tyler, et al.
Published: (2026) -
Practical challenges of control monitoring in frontier AI deployments
by: Lindner, David, et al.
Published: (2025) -
LLMs + Security = Trouble
by: Livshits, Benjamin
Published: (2026)