Containment Verification: AI Safety Guarantees Independent of Alignment
Fuente:
arXiv
Saved in:
| Main Authors: | Moon, Royce, Varshney, Lav R. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Symbolic Guardrails for Domain-Specific Agents: Stronger Safety and Security Guarantees Without Sacrificing Utility
by: Hong, Yining, et al.
Published: (2026)
by: Hong, Yining, et al.
Published: (2026)
Broken by Default: A Formal Verification Study of Security Vulnerabilities in AI-Generated Code
by: Blain, Dominik, et al.
Published: (2026)
by: Blain, Dominik, et al.
Published: (2026)
Natural Language based Specification and Verification
by: Li, Zhaorui, et al.
Published: (2026)
by: Li, Zhaorui, et al.
Published: (2026)
Safety and Performance, Why Not Both? Bi-Objective Optimized Model Compression against Heterogeneous Attacks Toward AI Software Deployment
by: Zhu, Jie, et al.
Published: (2024)
by: Zhu, Jie, et al.
Published: (2024)
Differentiated Directional Intervention A Framework for Evading LLM Safety Alignment
by: Zhang, Peng, et al.
Published: (2025)
by: Zhang, Peng, et al.
Published: (2025)
Inverting the Shield: Systematically Generating Safety Tests from Policy Specifications
by: Lu, Xiaoyue, et al.
Published: (2026)
by: Lu, Xiaoyue, et al.
Published: (2026)
AdaptiveGuard: Towards Adaptive Runtime Safety for LLM-Powered Software
by: Yang, Rui, et al.
Published: (2025)
by: Yang, Rui, et al.
Published: (2025)
AI security and cyber risk in IoT systems
by: Radanliev, Petar, et al.
Published: (2024)
by: Radanliev, Petar, et al.
Published: (2024)
Benchmark of Benchmarks: Unpacking Influence and Code Repository Quality in LLM Safety Benchmarks
by: Chu, Junjie, et al.
Published: (2026)
by: Chu, Junjie, et al.
Published: (2026)
Securing the Future of IVR: AI-Driven Innovation with Agile Security, Data Regulation, and Ethical AI Integration
by: Shaikh, Khushbu Mehboob, et al.
Published: (2025)
by: Shaikh, Khushbu Mehboob, et al.
Published: (2025)
A sketch of an AI control safety case
by: Korbak, Tomek, et al.
Published: (2025)
by: Korbak, Tomek, et al.
Published: (2025)
AI Code Generators for Security: Friend or Foe?
by: Natella, Roberto, et al.
Published: (2024)
by: Natella, Roberto, et al.
Published: (2024)
Risks of ignoring uncertainty propagation in AI-augmented security pipelines
by: Mezzi, Emanuele, et al.
Published: (2024)
by: Mezzi, Emanuele, et al.
Published: (2024)
Identifying the Supply Chain of AI for Trustworthiness and Risk Management in Critical Applications
by: Sheh, Raymond K., et al.
Published: (2025)
by: Sheh, Raymond K., et al.
Published: (2025)
DevOps-Gym: Benchmarking AI Agents in Software DevOps Cycle
by: Tang, Yuheng, et al.
Published: (2026)
by: Tang, Yuheng, et al.
Published: (2026)
Testing Storage-System Correctness: Challenges, Fuzzing Limitations, and AI-Augmented Opportunities
by: Wang, Ying, et al.
Published: (2026)
by: Wang, Ying, et al.
Published: (2026)
Poisoning Programs by Un-Repairing Code: Security Concerns of AI-generated Code
by: Improta, Cristina
Published: (2024)
by: Improta, Cristina
Published: (2024)
Data and Context Matter: Towards Generalizing AI-based Software Vulnerability Detection
by: Safdar, Rijha, et al.
Published: (2025)
by: Safdar, Rijha, et al.
Published: (2025)
Block MedCare: Advancing healthcare through blockchain integration with AI and IoT
by: Simonoski, Oliver, et al.
Published: (2024)
by: Simonoski, Oliver, et al.
Published: (2024)
AIBoMGen: Generating an AI Bill of Materials for Secure, Transparent, and Compliant Model Training
by: Vandendriessche, Wiebe, et al.
Published: (2026)
by: Vandendriessche, Wiebe, et al.
Published: (2026)
Zer0n: An AI-Assisted Vulnerability Discovery and Blockchain-Backed Integrity Framework
by: Parmar, Harshil, et al.
Published: (2026)
by: Parmar, Harshil, et al.
Published: (2026)
LinuxArena: A Control Setting for AI Agents in Live Production Software Environments
by: Tracy, Tyler, et al.
Published: (2026)
by: Tracy, Tyler, et al.
Published: (2026)
Investigating Detection and Obfuscation of Prompt Injection Attacks Against Software Reverse Engineering AI Agents
by: Crawford, Brian, et al.
Published: (2026)
by: Crawford, Brian, et al.
Published: (2026)
SOK: Exploring Hallucinations and Security Risks in AI-Assisted Software Development with Insights for LLM Deployment
by: Haque, Ariful, et al.
Published: (2025)
by: Haque, Ariful, et al.
Published: (2025)
MalCodeAI: Autonomous Vulnerability Detection and Remediation via Language Agnostic Code Reasoning
by: Gajjar, Jugal, et al.
Published: (2025)
by: Gajjar, Jugal, et al.
Published: (2025)
LLMSecConfig: An LLM-Based Approach for Fixing Software Container Misconfigurations
by: Ye, Ziyang, et al.
Published: (2025)
by: Ye, Ziyang, et al.
Published: (2025)
Rethinking Autonomy: Preventing Failures in AI-Driven Software Engineering
by: Navneet, Satyam Kumar, et al.
Published: (2025)
by: Navneet, Satyam Kumar, et al.
Published: (2025)
Inference-Time Safety For Code LLMs Via Retrieval-Augmented Revision
by: Mukherjee, Manisha, et al.
Published: (2026)
by: Mukherjee, Manisha, et al.
Published: (2026)
Learning to Generate Secure Code via Token-Level Rewards
by: Quan, Jiazheng, et al.
Published: (2026)
by: Quan, Jiazheng, et al.
Published: (2026)
MOSAIC-Bench: Measuring Compositional Vulnerability Induction in Coding Agents
by: Steinberg, Jonathan, et al.
Published: (2026)
by: Steinberg, Jonathan, et al.
Published: (2026)
TEMPLATEFUZZ: Fine-Grained Chat Template Fuzzing for Jailbreaking and Red Teaming LLMs
by: Shen, Qingchao, et al.
Published: (2026)
by: Shen, Qingchao, et al.
Published: (2026)
When Safe Skills Collide: Measuring Compositional Risk in Agent Skill Ecosystems
by: Wang, Su, et al.
Published: (2026)
by: Wang, Su, et al.
Published: (2026)
Favia: Forensic Agent for Vulnerability-fix Identification and Analysis
by: Storhaug, André, et al.
Published: (2026)
by: Storhaug, André, et al.
Published: (2026)
Beyond BeautifulSoup: Benchmarking LLM-Powered Web Scraping for Everyday Users
by: Bhardwaj, Arth, et al.
Published: (2026)
by: Bhardwaj, Arth, et al.
Published: (2026)
OpenSage: Self-programming Agent Generation Engine
by: Li, Hongwei, et al.
Published: (2026)
by: Li, Hongwei, et al.
Published: (2026)
SecCodeBench-V2 Technical Report
by: Chen, Longfei, et al.
Published: (2026)
by: Chen, Longfei, et al.
Published: (2026)
Security of LLM-generated Code: A Comparative Analysis
by: Morkonda, Srivathsan G, et al.
Published: (2026)
by: Morkonda, Srivathsan G, et al.
Published: (2026)
SCDBench: A Benchmark for LLM-Based Smart Contract Decompilers
by: Qin, Kaihua, et al.
Published: (2026)
by: Qin, Kaihua, et al.
Published: (2026)
Decaf: Improving Neural Decompilation with Automatic Feedback and Search
by: Shypula, Alexander, et al.
Published: (2026)
by: Shypula, Alexander, et al.
Published: (2026)
From SFT to RL: Demystifying the Post-Training Pipeline for LLM-based Vulnerability Detection
by: Li, Youpeng, et al.
Published: (2026)
by: Li, Youpeng, et al.
Published: (2026)
Similar Items
-
Symbolic Guardrails for Domain-Specific Agents: Stronger Safety and Security Guarantees Without Sacrificing Utility
by: Hong, Yining, et al.
Published: (2026) -
Broken by Default: A Formal Verification Study of Security Vulnerabilities in AI-Generated Code
by: Blain, Dominik, et al.
Published: (2026) -
Natural Language based Specification and Verification
by: Li, Zhaorui, et al.
Published: (2026) -
Safety and Performance, Why Not Both? Bi-Objective Optimized Model Compression against Heterogeneous Attacks Toward AI Software Deployment
by: Zhu, Jie, et al.
Published: (2024) -
Differentiated Directional Intervention A Framework for Evading LLM Safety Alignment
by: Zhang, Peng, et al.
Published: (2025)