Solver-Aided Verification of Policy Compliance in Tool-Augmented LLM Agents
Fuente:
arXiv
Guardado en:
| Autores principales: | Winston, Cailin, Winston, Claris, Just, René |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
ARPaCCino: An Agentic-RAG for Policy as Code Compliance
por: Romeo, Francesco, et al.
Publicado: (2025)
por: Romeo, Francesco, et al.
Publicado: (2025)
MCP-Zero: Active Tool Discovery for Autonomous LLM Agents
por: Fei, Xiang, et al.
Publicado: (2025)
por: Fei, Xiang, et al.
Publicado: (2025)
PropertyGPT: LLM-driven Formal Verification of Smart Contracts through Retrieval-Augmented Property Generation
por: Liu, Ye, et al.
Publicado: (2024)
por: Liu, Ye, et al.
Publicado: (2024)
CircuChain: Disentangling Competence and Compliance in LLM Circuit Analysis
por: Ravishankara, Mayank
Publicado: (2026)
por: Ravishankara, Mayank
Publicado: (2026)
AgentGuard: Runtime Verification of AI Agents
por: Koohestani, Roham
Publicado: (2025)
por: Koohestani, Roham
Publicado: (2025)
Investigating Tool-Memory Conflicts in Tool-Augmented LLMs
por: Cheng, Jiali, et al.
Publicado: (2026)
por: Cheng, Jiali, et al.
Publicado: (2026)
Schema First Tool APIs for LLM Agents: A Controlled Study of Tool Misuse, Recovery, and Budgeted Performance
por: Sigdel, Akshey, et al.
Publicado: (2026)
por: Sigdel, Akshey, et al.
Publicado: (2026)
VerilogReader: LLM-Aided Hardware Test Generation
por: Ma, Ruiyang, et al.
Publicado: (2024)
por: Ma, Ruiyang, et al.
Publicado: (2024)
Adopting RAG for LLM-Aided Future Vehicle Design
por: Zolfaghari, Vahid, et al.
Publicado: (2024)
por: Zolfaghari, Vahid, et al.
Publicado: (2024)
RAG-MCP: Mitigating Prompt Bloat in LLM Tool Selection via Retrieval-Augmented Generation
por: Gan, Tiantian, et al.
Publicado: (2025)
por: Gan, Tiantian, et al.
Publicado: (2025)
Graph-Based Self-Healing Tool Routing for Cost-Efficient LLM Agents
por: Bholani, Neeraj
Publicado: (2026)
por: Bholani, Neeraj
Publicado: (2026)
ComplexMCP: Evaluation of LLM Agents in Dynamic, Interdependent, and Large-Scale Tool Sandbox
por: Li, Yuanyang, et al.
Publicado: (2026)
por: Li, Yuanyang, et al.
Publicado: (2026)
ToolFuzz -- Automated Agent Tool Testing
por: Milev, Ivan, et al.
Publicado: (2025)
por: Milev, Ivan, et al.
Publicado: (2025)
Z-Space: A Multi-Agent Tool Orchestration Framework for Enterprise-Grade LLM Automation
por: He, Qingsong, et al.
Publicado: (2025)
por: He, Qingsong, et al.
Publicado: (2025)
Verification-Guided Context Optimization for Tool Calling via Hierarchical LLMs-as-Editors
por: Li, Henger, et al.
Publicado: (2025)
por: Li, Henger, et al.
Publicado: (2025)
Automating Android Build Repair: Bridging the Reasoning-Execution Gap in LLM Agents with Domain-Specific Tools
por: Son, Ha Min, et al.
Publicado: (2025)
por: Son, Ha Min, et al.
Publicado: (2025)
Butterfly Effects in Toolchains: A Comprehensive Analysis of Failed Parameter Filling in LLM Tool-Agent Systems
por: Xiong, Qian, et al.
Publicado: (2025)
por: Xiong, Qian, et al.
Publicado: (2025)
Breaking the Illusion of Identity in LLM Tooling
por: Miller, Marek
Publicado: (2026)
por: Miller, Marek
Publicado: (2026)
ToolPRMBench: Evaluating and Advancing Process Reward Models for Tool-using Agents
por: Li, Dawei, et al.
Publicado: (2026)
por: Li, Dawei, et al.
Publicado: (2026)
Evaluating Plan Compliance in Autonomous Programming Agents
por: Liu, Shuyang, et al.
Publicado: (2026)
por: Liu, Shuyang, et al.
Publicado: (2026)
RAG-Verus: Repository-Level Program Verification with LLMs using Retrieval Augmented Generation
por: Zhong, Sicheng, et al.
Publicado: (2025)
por: Zhong, Sicheng, et al.
Publicado: (2025)
LLM Company Policies and Policy Implications in Software Organizations
por: Khojah, Ranim, et al.
Publicado: (2025)
por: Khojah, Ranim, et al.
Publicado: (2025)
Does Pass Rate Tell the Whole Story? Evaluating Design Constraint Compliance in LLM-based Issue Resolution
por: Yu, Kai, et al.
Publicado: (2026)
por: Yu, Kai, et al.
Publicado: (2026)
Verifying LLM-Generated Code in the Context of Software Verification with Ada/SPARK
por: Cramer, Marcos, et al.
Publicado: (2025)
por: Cramer, Marcos, et al.
Publicado: (2025)
EvolveTool-Bench: Evaluating the Quality of LLM-Generated Tool Libraries as Software Artifacts
por: Kaliyev, Alibek T., et al.
Publicado: (2026)
por: Kaliyev, Alibek T., et al.
Publicado: (2026)
The Tool-Overuse Illusion: Why Does LLM Prefer External Tools over Internal Knowledge?
por: Zeng, Yirong, et al.
Publicado: (2026)
por: Zeng, Yirong, et al.
Publicado: (2026)
The Causal Impact of Tool Affordance on Safety Alignment in LLM Agents
por: Yu, Shasha, et al.
Publicado: (2026)
por: Yu, Shasha, et al.
Publicado: (2026)
Vision2Web: A Hierarchical Benchmark for Visual Website Development with Agent Verification
por: He, Zehai, et al.
Publicado: (2026)
por: He, Zehai, et al.
Publicado: (2026)
AISysRev -- LLM-based Tool for Title-abstract Screening
por: Huotala, Aleksi, et al.
Publicado: (2025)
por: Huotala, Aleksi, et al.
Publicado: (2025)
ASA: Training-Free Representation Engineering for Tool-Calling Agents
por: Wang, Youjin, et al.
Publicado: (2026)
por: Wang, Youjin, et al.
Publicado: (2026)
Squeez: Task-Conditioned Tool-Output Pruning for Coding Agents
por: Kovács, Ádám
Publicado: (2026)
por: Kovács, Ádám
Publicado: (2026)
UnitTenX: Generating Tests for Legacy Packages with AI Agents Powered by Formal Verification
por: Charalambous, Yiannis, et al.
Publicado: (2025)
por: Charalambous, Yiannis, et al.
Publicado: (2025)
RepairAgent: An Autonomous, LLM-Based Agent for Program Repair
por: Bouzenia, Islem, et al.
Publicado: (2024)
por: Bouzenia, Islem, et al.
Publicado: (2024)
Automatic Red Teaming LLM-based Agents with Model Context Protocol Tools
por: He, Ping, et al.
Publicado: (2025)
por: He, Ping, et al.
Publicado: (2025)
Enhancing Legal Compliance and Regulation Analysis with Large Language Models
por: Hassani, Shabnam
Publicado: (2024)
por: Hassani, Shabnam
Publicado: (2024)
Verification Limits Code LLM Training
por: Gureja, Srishti, et al.
Publicado: (2025)
por: Gureja, Srishti, et al.
Publicado: (2025)
JTPRO: A Joint Tool-Prompt Reflective Optimization Framework for Language Agents
por: Ghoshal, Sandip, et al.
Publicado: (2026)
por: Ghoshal, Sandip, et al.
Publicado: (2026)
SAGE: Tool-Augmented LLM Task Solving Strategies in Scalable Multi-Agent Environments
por: Strehlow, Robert K., et al.
Publicado: (2026)
por: Strehlow, Robert K., et al.
Publicado: (2026)
Reducing Cost of LLM Agents with Trajectory Reduction
por: Xiao, Yuan-An, et al.
Publicado: (2025)
por: Xiao, Yuan-An, et al.
Publicado: (2025)
LLM Collaboration With Multi-Agent Reinforcement Learning
por: Liu, Shuo, et al.
Publicado: (2025)
por: Liu, Shuo, et al.
Publicado: (2025)
Ejemplares similares
-
ARPaCCino: An Agentic-RAG for Policy as Code Compliance
por: Romeo, Francesco, et al.
Publicado: (2025) -
MCP-Zero: Active Tool Discovery for Autonomous LLM Agents
por: Fei, Xiang, et al.
Publicado: (2025) -
PropertyGPT: LLM-driven Formal Verification of Smart Contracts through Retrieval-Augmented Property Generation
por: Liu, Ye, et al.
Publicado: (2024) -
CircuChain: Disentangling Competence and Compliance in LLM Circuit Analysis
por: Ravishankara, Mayank
Publicado: (2026) -
AgentGuard: Runtime Verification of AI Agents
por: Koohestani, Roham
Publicado: (2025)