VeriGuard: Enhancing LLM Agent Safety via Verified Code Generation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Miculicich, Lesly, Parmar, Mihir, Palangi, Hamid, Dvijotham, Krishnamurthy Dj, Montanari, Mirko, Pfister, Tomas, Le, Long T.
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914078792876032
author Miculicich, Lesly
Parmar, Mihir
Palangi, Hamid
Dvijotham, Krishnamurthy Dj
Montanari, Mirko
Pfister, Tomas
Le, Long T.
author_facet Miculicich, Lesly
Parmar, Mihir
Palangi, Hamid
Dvijotham, Krishnamurthy Dj
Montanari, Mirko
Pfister, Tomas
Le, Long T.
contents The deployment of autonomous AI agents in sensitive domains, such as healthcare, introduces critical risks to safety, security, and privacy. These agents may deviate from user objectives, violate data handling policies, or be compromised by adversarial attacks. Mitigating these dangers necessitates a mechanism to formally guarantee that an agent's actions adhere to predefined safety constraints, a challenge that existing systems do not fully address. We introduce VeriGuard, a novel framework that provides formal safety guarantees for LLM-based agents through a dual-stage architecture designed for robust and verifiable correctness. The initial offline stage involves a comprehensive validation process. It begins by clarifying user intent to establish precise safety specifications. VeriGuard then synthesizes a behavioral policy and subjects it to both testing and formal verification to prove its compliance with these specifications. This iterative process refines the policy until it is deemed correct. Subsequently, the second stage provides online action monitoring, where VeriGuard operates as a runtime monitor to validate each proposed agent action against the pre-verified policy before execution. This separation of the exhaustive offline validation from the lightweight online monitoring allows formal guarantees to be practically applied, providing a robust safeguard that substantially improves the trustworthiness of LLM agents.
format Preprint
id arxiv_https___arxiv_org_abs_2510_05156
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VeriGuard: Enhancing LLM Agent Safety via Verified Code Generation
Miculicich, Lesly
Parmar, Mihir
Palangi, Hamid
Dvijotham, Krishnamurthy Dj
Montanari, Mirko
Pfister, Tomas
Le, Long T.
Software Engineering
Artificial Intelligence
Cryptography and Security
I.2.7
The deployment of autonomous AI agents in sensitive domains, such as healthcare, introduces critical risks to safety, security, and privacy. These agents may deviate from user objectives, violate data handling policies, or be compromised by adversarial attacks. Mitigating these dangers necessitates a mechanism to formally guarantee that an agent's actions adhere to predefined safety constraints, a challenge that existing systems do not fully address. We introduce VeriGuard, a novel framework that provides formal safety guarantees for LLM-based agents through a dual-stage architecture designed for robust and verifiable correctness. The initial offline stage involves a comprehensive validation process. It begins by clarifying user intent to establish precise safety specifications. VeriGuard then synthesizes a behavioral policy and subjects it to both testing and formal verification to prove its compliance with these specifications. This iterative process refines the policy until it is deemed correct. Subsequently, the second stage provides online action monitoring, where VeriGuard operates as a runtime monitor to validate each proposed agent action against the pre-verified policy before execution. This separation of the exhaustive offline validation from the lightweight online monitoring allows formal guarantees to be practically applied, providing a robust safeguard that substantially improves the trustworthiness of LLM agents.
title VeriGuard: Enhancing LLM Agent Safety via Verified Code Generation
topic Software Engineering
Artificial Intelligence
Cryptography and Security
I.2.7
url https://arxiv.org/abs/2510.05156