ToolSafe: Enhancing Tool Invocation Safety of LLM-based agents via Proactive Step-level Guardrail and Feedback
Fuente:
arXiv
Guardado en:
| Autores principales: | Mou, Yutao, Xue, Zhangchi, Li, Lijun, Liu, Peiyang, Zhang, Shikun, Ye, Wei, Shao, Jing |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
SaRO: Enhancing LLM Safety through Reasoning-based Alignment
por: Mou, Yutao, et al.
Publicado: (2025)
por: Mou, Yutao, et al.
Publicado: (2025)
SG-Bench: Evaluating LLM Safety Generalization Across Diverse Tasks and Prompt Types
por: Mou, Yutao, et al.
Publicado: (2024)
por: Mou, Yutao, et al.
Publicado: (2024)
TraceSafe: A Systematic Assessment of LLM Guardrails on Multi-Step Tool-Calling Trajectories
por: Chen, Yen-Shan, et al.
Publicado: (2026)
por: Chen, Yen-Shan, et al.
Publicado: (2026)
CodeTool: Enhancing Programmatic Tool Invocation of LLMs via Process Supervision
por: Lu, Yifei, et al.
Publicado: (2025)
por: Lu, Yifei, et al.
Publicado: (2025)
Decoupling Safety into Orthogonal Subspace: Cost-Efficient and Performance-Preserving Alignment for Large Language Models
por: Mou, Yutao, et al.
Publicado: (2025)
por: Mou, Yutao, et al.
Publicado: (2025)
Autoformalizer with Tool Feedback
por: Guo, Qi, et al.
Publicado: (2025)
por: Guo, Qi, et al.
Publicado: (2025)
Building Effective Safety Guardrails in AI Education Tools
por: Clark, Hannah-Beth, et al.
Publicado: (2025)
por: Clark, Hannah-Beth, et al.
Publicado: (2025)
Re-Invoke: Tool Invocation Rewriting for Zero-Shot Tool Retrieval
por: Chen, Yanfei, et al.
Publicado: (2024)
por: Chen, Yanfei, et al.
Publicado: (2024)
Advancing and Benchmarking Personalized Tool Invocation for LLMs
por: Huang, Xu, et al.
Publicado: (2025)
por: Huang, Xu, et al.
Publicado: (2025)
OSWorld-MCP: Benchmarking MCP Tool Invocation In Computer-Use Agents
por: Jia, Hongrui, et al.
Publicado: (2025)
por: Jia, Hongrui, et al.
Publicado: (2025)
AttriGuard: Defeating Indirect Prompt Injection in LLM Agents via Causal Attribution of Tool Invocations
por: He, Yu, et al.
Publicado: (2026)
por: He, Yu, et al.
Publicado: (2026)
Divide-Then-Aggregate: An Efficient Tool Learning Method via Parallel Tool Invocation
por: Zhu, Dongsheng, et al.
Publicado: (2025)
por: Zhu, Dongsheng, et al.
Publicado: (2025)
Implicit Hierarchical GRPO: Decoupling Tool Invocation from Execution for Tool-Integrated Mathematical Reasoning
por: Wang, Li, et al.
Publicado: (2026)
por: Wang, Li, et al.
Publicado: (2026)
PsySafe: A Comprehensive Framework for Psychological-based Attack, Defense, and Evaluation of Multi-agent System Safety
por: Zhang, Zaibin, et al.
Publicado: (2024)
por: Zhang, Zaibin, et al.
Publicado: (2024)
Can You Really Trust Code Copilots? Evaluating Large Language Models from a Code Security Perspective
por: Mou, Yutao, et al.
Publicado: (2025)
por: Mou, Yutao, et al.
Publicado: (2025)
Do LLMs Know Tool Irrelevance? Demystifying Structural Alignment Bias in Tool Invocations
por: Liu, Yilong, et al.
Publicado: (2026)
por: Liu, Yilong, et al.
Publicado: (2026)
Evaluating Implicit Regulatory Compliance in LLM Tool Invocation via Logic-Guided Synthesis
por: Song, Da, et al.
Publicado: (2026)
por: Song, Da, et al.
Publicado: (2026)
SafeHarbor: Hierarchical Memory-Augmented Guardrail for LLM Agent Safety
por: Liu, Zhe, et al.
Publicado: (2026)
por: Liu, Zhe, et al.
Publicado: (2026)
SafeVision: Efficient Image Guardrail with Robust Policy Adherence and Explainability
por: Xu, Peiyang, et al.
Publicado: (2025)
por: Xu, Peiyang, et al.
Publicado: (2025)
Automated Creation and Enrichment Framework for Improved Invocation of Enterprise APIs as Tools
por: Agarwal, Prerna, et al.
Publicado: (2025)
por: Agarwal, Prerna, et al.
Publicado: (2025)
GeoMind: An Agentic Workflow for Lithology Classification with Reasoned Tool Invocation
por: Zhou, Yitong, et al.
Publicado: (2026)
por: Zhou, Yitong, et al.
Publicado: (2026)
VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation
por: Zhao, Yiming, et al.
Publicado: (2026)
por: Zhao, Yiming, et al.
Publicado: (2026)
Break the Optimization Barrier of LLM-Enhanced Recommenders: A Theoretical Analysis and Practical Framework
por: Zhu, Zhangchi, et al.
Publicado: (2026)
por: Zhu, Zhangchi, et al.
Publicado: (2026)
When Agents Fail to Act: A Diagnostic Framework for Tool Invocation Reliability in Multi-Agent LLM Systems
por: Huang, Donghao, et al.
Publicado: (2026)
por: Huang, Donghao, et al.
Publicado: (2026)
Safety Guardrails for LLM-Enabled Robots
por: Ravichandran, Zachary, et al.
Publicado: (2025)
por: Ravichandran, Zachary, et al.
Publicado: (2025)
StepTool: Enhancing Multi-Step Tool Usage in LLMs via Step-Grained Reinforcement Learning
por: Yu, Yuanqing, et al.
Publicado: (2024)
por: Yu, Yuanqing, et al.
Publicado: (2024)
AgentBelt: Runtime Guardrails for LLM Agent Tool Calls — ASE 2026 Artifact
por: Anonymous
Publicado: (2026)
por: Anonymous
Publicado: (2026)
PSG-Agent: Personality-Aware Safety Guardrail for LLM-based Agents
por: Wu, Yaozu, et al.
Publicado: (2025)
por: Wu, Yaozu, et al.
Publicado: (2025)
Enhancing Guardrails for Safe and Secure Healthcare AI
por: Gangavarapu, Ananya
Publicado: (2024)
por: Gangavarapu, Ananya
Publicado: (2024)
ToolPlanner: A Tool Augmented LLM for Multi Granularity Instructions with Path Planning and Feedback
por: Wu, Qinzhuo, et al.
Publicado: (2024)
por: Wu, Qinzhuo, et al.
Publicado: (2024)
ConsisGuard: Aligning Safety Deliberation with Policy Enforcement in LLM Guardrails
por: Wang, Yan, et al.
Publicado: (2026)
por: Wang, Yan, et al.
Publicado: (2026)
SafeToolBench: Pioneering a Prospective Benchmark to Evaluating Tool Utilization Safety in LLMs
por: Xia, Hongfei, et al.
Publicado: (2025)
por: Xia, Hongfei, et al.
Publicado: (2025)
Trust No Tool: Evaluating and Defending LLM Agents under Untrusted Tool Feedback
por: Yan, Lecheng, et al.
Publicado: (2026)
por: Yan, Lecheng, et al.
Publicado: (2026)
RSAgent: Learning to Reason and Act for Text-Guided Segmentation via Multi-Turn Tool Invocations
por: He, Xingqi, et al.
Publicado: (2025)
por: He, Xingqi, et al.
Publicado: (2025)
Red-Teaming Coding Agents from a Tool-Invocation Perspective: An Empirical Security Assessment
por: Xie, Yuchong, et al.
Publicado: (2025)
por: Xie, Yuchong, et al.
Publicado: (2025)
Guardrails as Infrastructure: Policy-First Control for Tool-Orchestrated Workflows
por: Sigdel, Akshey, et al.
Publicado: (2026)
por: Sigdel, Akshey, et al.
Publicado: (2026)
SEARL: Joint Optimization of Policy and Tool Graph Memory for Self-Evolving Agents
por: Feng, Xinshun, et al.
Publicado: (2026)
por: Feng, Xinshun, et al.
Publicado: (2026)
Towards Verifiably Safe Tool Use for LLM Agents
por: Doshi, Aarya, et al.
Publicado: (2026)
por: Doshi, Aarya, et al.
Publicado: (2026)
ProGuard: Towards Proactive Multimodal Safeguard
por: Yu, Shaohan, et al.
Publicado: (2025)
por: Yu, Shaohan, et al.
Publicado: (2025)
HGMF: A Hierarchical Gaussian Mixture Framework for Scalable Tool Invocation within the Model Context Protocol
por: Xing, Wenpeng, et al.
Publicado: (2025)
por: Xing, Wenpeng, et al.
Publicado: (2025)
Ejemplares similares
-
SaRO: Enhancing LLM Safety through Reasoning-based Alignment
por: Mou, Yutao, et al.
Publicado: (2025) -
SG-Bench: Evaluating LLM Safety Generalization Across Diverse Tasks and Prompt Types
por: Mou, Yutao, et al.
Publicado: (2024) -
TraceSafe: A Systematic Assessment of LLM Guardrails on Multi-Step Tool-Calling Trajectories
por: Chen, Yen-Shan, et al.
Publicado: (2026) -
CodeTool: Enhancing Programmatic Tool Invocation of LLMs via Process Supervision
por: Lu, Yifei, et al.
Publicado: (2025) -
Decoupling Safety into Orthogonal Subspace: Cost-Efficient and Performance-Preserving Alignment for Large Language Models
por: Mou, Yutao, et al.
Publicado: (2025)