VeriGuard: Enhancing LLM Agent Safety via Verified Code Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Miculicich, Lesly, Parmar, Mihir, Palangi, Hamid, Dvijotham, Krishnamurthy Dj, Montanari, Mirko, Pfister, Tomas, Le, Long T. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CausalArmor: Efficient Indirect Prompt Injection Guardrails via Causal Attribution
by: Kim, Minbeom, et al.
Published: (2026)
by: Kim, Minbeom, et al.
Published: (2026)
Mind the GAP: Text Safety Does Not Transfer to Tool-Call Safety in LLM Agents
by: Cartagena, Arnold, et al.
Published: (2026)
by: Cartagena, Arnold, et al.
Published: (2026)
Synapse: Adaptive Arbitration of Complementary Expertise in Time Series Foundational Models
by: Das, Sarkar Snigdha Sarathi, et al.
Published: (2025)
by: Das, Sarkar Snigdha Sarathi, et al.
Published: (2025)
Agentic Refactoring: An Empirical Study of AI Coding Agents
by: Horikawa, Kosei, et al.
Published: (2025)
by: Horikawa, Kosei, et al.
Published: (2025)
LiSA: Lifelong Safety Adaptation via Conservative Policy Induction
by: Kim, Minbeom, et al.
Published: (2026)
by: Kim, Minbeom, et al.
Published: (2026)
ScholarPeer: A Context-Aware Multi-Agent Framework for Automated Peer Review
by: Goyal, Palash, et al.
Published: (2026)
by: Goyal, Palash, et al.
Published: (2026)
Toward Architecture-Aware Evaluation Metrics for LLM Agents
by: Souza, Débora, et al.
Published: (2026)
by: Souza, Débora, et al.
Published: (2026)
When Code Smells Meet ML: On the Lifecycle of ML-specific Code Smells in ML-enabled Systems
by: Recupito, Gilberto, et al.
Published: (2024)
by: Recupito, Gilberto, et al.
Published: (2024)
Energy-Aware Decision Making in Software Stack Upgrades
by: Stocker, Mirko, et al.
Published: (2026)
by: Stocker, Mirko, et al.
Published: (2026)
Prompt Engineering Strategies for LLM-based Qualitative Coding of Psychological Safety in Software Engineering Communities: A Controlled Empirical Study
by: Alshaikh, Moaath, et al.
Published: (2026)
by: Alshaikh, Moaath, et al.
Published: (2026)
LLM-Based Multi-Agent Blackboard System for Information Discovery in Data Science
by: Salemi, Alireza, et al.
Published: (2025)
by: Salemi, Alireza, et al.
Published: (2025)
Developer-LLM Conversations: An Empirical Study of Interactions and Generated Code Quality
by: Zhong, Suzhen, et al.
Published: (2025)
by: Zhong, Suzhen, et al.
Published: (2025)
ProbGuard: Probabilistic Runtime Monitoring for LLM Agent Safety
by: Wang, Haoyu, et al.
Published: (2025)
by: Wang, Haoyu, et al.
Published: (2025)
AdaDec: A Uncertainty-Guided Lookahead Decoding Framework for LLM-Based Code Generation
by: He, Kaifeng, et al.
Published: (2025)
by: He, Kaifeng, et al.
Published: (2025)
LLM-based vs. Search-based Merge Conflict Resolution: An Empirical Study of Competing Paradigms
by: Junior, Heleno de Souza Campos, et al.
Published: (2026)
by: Junior, Heleno de Souza Campos, et al.
Published: (2026)
CaLM: Contrasting Large and Small Language Models to Verify Grounded Generation
by: Hsu, I-Hung, et al.
Published: (2024)
by: Hsu, I-Hung, et al.
Published: (2024)
LLM Security Guard for Code
by: Kavian, Arya, et al.
Published: (2024)
by: Kavian, Arya, et al.
Published: (2024)
Learning Software Bug Reports: A Systematic Literature Review
by: Long, Guoming, et al.
Published: (2025)
by: Long, Guoming, et al.
Published: (2025)
Lore: Repurposing Git Commit Messages as a Structured Knowledge Protocol for AI Coding Agents
by: Stetsenko, Ivan
Published: (2026)
by: Stetsenko, Ivan
Published: (2026)
ReDel: A Toolkit for LLM-Powered Recursive Multi-Agent Systems
by: Zhu, Andrew, et al.
Published: (2024)
by: Zhu, Andrew, et al.
Published: (2024)
LLMs as Idiomatic Decompilers: Recovering High-Level Code from x86-64 Assembly for Dart
by: Abualazm, Raafat, et al.
Published: (2026)
by: Abualazm, Raafat, et al.
Published: (2026)
ContractBench: Can LLM Agents Preserve Observation Contracts?
by: Wang, Jicheng, et al.
Published: (2026)
by: Wang, Jicheng, et al.
Published: (2026)
Leveraging Test Driven Development with Large Language Models for Reliable and Verifiable Spreadsheet Code Generation: A Research Framework
by: Thorne, Simon, et al.
Published: (2025)
by: Thorne, Simon, et al.
Published: (2025)
XPath Agent: An Efficient XPath Programming Agent Based on LLM for Web Crawler
by: Li, Yu, et al.
Published: (2024)
by: Li, Yu, et al.
Published: (2024)
Automated Code Review Using Large Language Models at Ericsson: An Experience Report
by: Ramesh, Shweta, et al.
Published: (2025)
by: Ramesh, Shweta, et al.
Published: (2025)
Enhancing LLM Code Generation Capabilities through Test-Driven Development and Code Interpreter
by: Jalil, Sajed, et al.
Published: (2025)
by: Jalil, Sajed, et al.
Published: (2025)
CWM: An Open-Weights LLM for Research on Code Generation with World Models
by: FAIR CodeGen team, et al.
Published: (2025)
by: FAIR CodeGen team, et al.
Published: (2025)
AgentOps: Enabling Observability of LLM Agents
by: Dong, Liming, et al.
Published: (2024)
by: Dong, Liming, et al.
Published: (2024)
A Framework for Testing and Adapting REST APIs as LLM Tools
by: Bandlamudi, Jayachandu, et al.
Published: (2025)
by: Bandlamudi, Jayachandu, et al.
Published: (2025)
The Upper Bound of Information Diffusion in Code Review
by: Dorner, Michael, et al.
Published: (2023)
by: Dorner, Michael, et al.
Published: (2023)
Comparative Analysis of LLM Abliteration Methods: A Cross-Architecture Evaluation
by: Young, Richard J.
Published: (2025)
by: Young, Richard J.
Published: (2025)
Towards Identifying Code Proficiency through the Analysis of Python Textbooks
by: Rojpaisarnkit, Ruksit, et al.
Published: (2024)
by: Rojpaisarnkit, Ruksit, et al.
Published: (2024)
Finetuning LLMs for Automatic Form Interaction on Web-Browser in Selenium Testing Framework
by: Le, Nguyen-Khang, et al.
Published: (2025)
by: Le, Nguyen-Khang, et al.
Published: (2025)
VeriScale: Adversarial Test-Suite Scaling for Verifiable Code Generation
by: Bai, Yifan, et al.
Published: (2026)
by: Bai, Yifan, et al.
Published: (2026)
VeriContest: A Competitive-Programming Benchmark for Verifiable Code Generation
by: Xie, Zichen, et al.
Published: (2026)
by: Xie, Zichen, et al.
Published: (2026)
AlgoVeri: An Aligned Benchmark for Verified Code Generation on Classical Algorithms
by: Zhao, Haoyu, et al.
Published: (2026)
by: Zhao, Haoyu, et al.
Published: (2026)
Safety Under Scaffolding: How Evaluation Conditions Shape Measured Safety
by: Gringras, David
Published: (2026)
by: Gringras, David
Published: (2026)
Mechanistic Understanding of Language Models in Syntactic Code Completion
by: Miller, Samuel, et al.
Published: (2025)
by: Miller, Samuel, et al.
Published: (2025)
CoSQA+: Pioneering the Multi-Choice Code Search Benchmark with Test-Driven Agents
by: Gong, Jing, et al.
Published: (2024)
by: Gong, Jing, et al.
Published: (2024)
Compiled AI: Deterministic Code Generation for LLM-Based Workflow Automation
by: Trooskens, Geert, et al.
Published: (2026)
by: Trooskens, Geert, et al.
Published: (2026)
Similar Items
-
CausalArmor: Efficient Indirect Prompt Injection Guardrails via Causal Attribution
by: Kim, Minbeom, et al.
Published: (2026) -
Mind the GAP: Text Safety Does Not Transfer to Tool-Call Safety in LLM Agents
by: Cartagena, Arnold, et al.
Published: (2026) -
Synapse: Adaptive Arbitration of Complementary Expertise in Time Series Foundational Models
by: Das, Sarkar Snigdha Sarathi, et al.
Published: (2025) -
Agentic Refactoring: An Empirical Study of AI Coding Agents
by: Horikawa, Kosei, et al.
Published: (2025) -
LiSA: Lifelong Safety Adaptation via Conservative Policy Induction
by: Kim, Minbeom, et al.
Published: (2026)