Towards Verifiably Safe Tool Use for LLM Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Doshi, Aarya, Hong, Yining, Xu, Congying, Kang, Eunsuk, Kapravelos, Alexandros, Kästner, Christian |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Symbolic Guardrails for Domain-Specific Agents: Stronger Safety and Security Guarantees Without Sacrificing Utility
by: Hong, Yining, et al.
Published: (2026)
by: Hong, Yining, et al.
Published: (2026)
FairSense: Long-Term Fairness Analysis of ML-Enabled Systems
by: She, Yining, et al.
Published: (2025)
by: She, Yining, et al.
Published: (2025)
Six Million (Suspected) Fake Stars in GitHub: A Growing Spiral of Popularity Contests, Spams, and Malware
by: He, Hao, et al.
Published: (2024)
by: He, Hao, et al.
Published: (2024)
From Hazard Identification to Controller Design: Proactive and LLM-Supported Safety Engineering for ML-Powered Systems
by: Hong, Yining, et al.
Published: (2025)
by: Hong, Yining, et al.
Published: (2025)
tl;dr: Chill, y'all: AI Will Not Devour SE
by: Kang, Eunsuk, et al.
Published: (2024)
by: Kang, Eunsuk, et al.
Published: (2024)
Counterexample Classification against Signal Temporal Logic Specifications
by: Zhang, Zhenya, et al.
Published: (2026)
by: Zhang, Zhenya, et al.
Published: (2026)
FASR: Automated Identification of Unsafe Control Actions in STPA
by: Dardik, Ian, et al.
Published: (2026)
by: Dardik, Ian, et al.
Published: (2026)
The Evolution of Tool Use in LLM Agents: From Single-Tool Call to Multi-Tool Orchestration
by: Xu, Haoyuan, et al.
Published: (2026)
by: Xu, Haoyuan, et al.
Published: (2026)
Multi-Agent Taint Specification Extraction for Vulnerability Detection
by: Ghebremichael, Jonah, et al.
Published: (2026)
by: Ghebremichael, Jonah, et al.
Published: (2026)
ToolScope: Enhancing LLM Agent Tool Use through Tool Merging and Context-Aware Filtering
by: Liu, Marianne Menglin, et al.
Published: (2025)
by: Liu, Marianne Menglin, et al.
Published: (2025)
SkillCraft: Can LLM Agents Learn to Use Tools Skillfully?
by: Chen, Shiqi, et al.
Published: (2026)
by: Chen, Shiqi, et al.
Published: (2026)
HarnessAgent: Scaling Automatic Fuzzing Harness Construction with Tool-Augmented LLM Pipelines
by: Yang, Kang, et al.
Published: (2025)
by: Yang, Kang, et al.
Published: (2025)
Towards Supporting Quality Architecture Evaluation with LLM Tools
by: Capilla, Rafael, et al.
Published: (2026)
by: Capilla, Rafael, et al.
Published: (2026)
OpenComputer: Verifiable Software Worlds for Computer-Use Agents
by: Wei, Jinbiao, et al.
Published: (2026)
by: Wei, Jinbiao, et al.
Published: (2026)
(Why) Is My Prompt Getting Worse? Rethinking Regression Testing for Evolving LLM APIs
by: Ma, Wanqin, et al.
Published: (2023)
by: Ma, Wanqin, et al.
Published: (2023)
Faver: Boosting LLM-based RTL Generation with Function Abstracted Verifiable Middleware
by: Mu, Jianan, et al.
Published: (2025)
by: Mu, Jianan, et al.
Published: (2025)
CompileAgent: Automated Real-World Repo-Level Compilation with Tool-Integrated LLM-based Agent System
by: Hu, Li, et al.
Published: (2025)
by: Hu, Li, et al.
Published: (2025)
CRScore++: Reinforcement Learning with Verifiable Tool and AI Feedback for Code Review
by: Kapadnis, Manav Nitin, et al.
Published: (2025)
by: Kapadnis, Manav Nitin, et al.
Published: (2025)
AgentSZZ: Teaching the LLM Agent to Play Detective with Bug-Inducing Commits
by: Lyu, Yunbo, et al.
Published: (2026)
by: Lyu, Yunbo, et al.
Published: (2026)
What Is Wrong with My Model? Identifying Systematic Problems with Semantic Data Slicing
by: Yang, Chenyang, et al.
Published: (2024)
by: Yang, Chenyang, et al.
Published: (2024)
Event-B Agent: Towards LLM Agent for Formal Model Synthesis and Repair
by: Wang, Hongshu, et al.
Published: (2026)
by: Wang, Hongshu, et al.
Published: (2026)
DADL: A Declarative Description Language for Enterprise Tool Libraries in LLM Agent Systems
by: Dunkel, Axel
Published: (2026)
by: Dunkel, Axel
Published: (2026)
User-Driven Adaptation: Tailoring Autonomous Driving Systems with Dynamic Preferences
by: Zhang, Mingyue, et al.
Published: (2024)
by: Zhang, Mingyue, et al.
Published: (2024)
Pretrained Embeddings as a Behavior Specification Mechanism
by: Kapoor, Parv, et al.
Published: (2025)
by: Kapoor, Parv, et al.
Published: (2025)
Gecko: A Simulation Environment with Stateful Feedback for Refining Agent Tool Calls
by: Zhang, Zeyu, et al.
Published: (2026)
by: Zhang, Zeyu, et al.
Published: (2026)
Teaching Code Refactoring Using LLMs
by: Khairnar, Anshul, et al.
Published: (2025)
by: Khairnar, Anshul, et al.
Published: (2025)
MR-Scout: Automated Synthesis of Metamorphic Relations from Existing Test Cases
by: Xu, Congying, et al.
Published: (2023)
by: Xu, Congying, et al.
Published: (2023)
VeriFix: Verifying Your Fix Towards An Atomicity Violation
by: Li, Zhuang, et al.
Published: (2025)
by: Li, Zhuang, et al.
Published: (2025)
Beyond pip install: Evaluating LLM Agents for the Automated Installation of Python Projects
by: Milliken, Louis, et al.
Published: (2024)
by: Milliken, Louis, et al.
Published: (2024)
How to Sustain a Scientific Open-Source Software Ecosystem: Learning from the Astropy Project
by: Sun, Jiayi, et al.
Published: (2024)
by: Sun, Jiayi, et al.
Published: (2024)
Does SWE-Bench-Verified Test Agent Ability or Model Memory?
by: Prathifkumar, Thanosan, et al.
Published: (2025)
by: Prathifkumar, Thanosan, et al.
Published: (2025)
MCP-Zero: Active Tool Discovery for Autonomous LLM Agents
by: Fei, Xiang, et al.
Published: (2025)
by: Fei, Xiang, et al.
Published: (2025)
Online-Optimized RAG for Tool Use and Function Calling
by: Pan, Yu, et al.
Published: (2025)
by: Pan, Yu, et al.
Published: (2025)
Can Emulating Semantic Translation Help LLMs with Code Translation? A Study Based on Pseudocode
by: Chen, Songqiang, et al.
Published: (2025)
by: Chen, Songqiang, et al.
Published: (2025)
Analyzing Maintenance Activities of Software Libraries
by: Tsakpinis, Alexandros
Published: (2023)
by: Tsakpinis, Alexandros
Published: (2023)
Towards Verified Code Reasoning by LLMs
by: Sistla, Meghana, et al.
Published: (2025)
by: Sistla, Meghana, et al.
Published: (2025)
AI Tool Use and Adoption in Software Development by Individuals and Organizations: A Grounded Theory Study
by: Li, Ze Shi, et al.
Published: (2024)
by: Li, Ze Shi, et al.
Published: (2024)
The Product Beyond the Model -- An Empirical Study of Repositories of Open-Source ML Products
by: Nahar, Nadia, et al.
Published: (2023)
by: Nahar, Nadia, et al.
Published: (2023)
Z-Space: A Multi-Agent Tool Orchestration Framework for Enterprise-Grade LLM Automation
by: He, Qingsong, et al.
Published: (2025)
by: He, Qingsong, et al.
Published: (2025)
Whose fault is it anyway? SILC: Safe Integration of LLM-Generated Code
by: Lin, Peisen, et al.
Published: (2024)
by: Lin, Peisen, et al.
Published: (2024)
Similar Items
-
Symbolic Guardrails for Domain-Specific Agents: Stronger Safety and Security Guarantees Without Sacrificing Utility
by: Hong, Yining, et al.
Published: (2026) -
FairSense: Long-Term Fairness Analysis of ML-Enabled Systems
by: She, Yining, et al.
Published: (2025) -
Six Million (Suspected) Fake Stars in GitHub: A Growing Spiral of Popularity Contests, Spams, and Malware
by: He, Hao, et al.
Published: (2024) -
From Hazard Identification to Controller Design: Proactive and LLM-Supported Safety Engineering for ML-Powered Systems
by: Hong, Yining, et al.
Published: (2025) -
tl;dr: Chill, y'all: AI Will Not Devour SE
by: Kang, Eunsuk, et al.
Published: (2024)