ToolFuzz -- Automated Agent Tool Testing
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Milev, Ivan, Balunović, Mislav, Baader, Maximilian, Vechev, Martin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Automated Benchmark Generation for Repository-Level Coding Tasks
von: Vergopoulos, Konstantinos, et al.
Veröffentlicht: (2025)
von: Vergopoulos, Konstantinos, et al.
Veröffentlicht: (2025)
AutoRestTest: A Tool for Automated REST API Testing Using LLMs and MARL
von: Stennett, Tyler, et al.
Veröffentlicht: (2025)
von: Stennett, Tyler, et al.
Veröffentlicht: (2025)
MIMIC-Py: An Extensible Tool for Personality-Driven Automated Game Testing with Large Language Models
von: Chen, Yifei, et al.
Veröffentlicht: (2026)
von: Chen, Yifei, et al.
Veröffentlicht: (2026)
SWT-Bench: Testing and Validating Real-World Bug-Fixes with Code Agents
von: Mündler, Niels, et al.
Veröffentlicht: (2024)
von: Mündler, Niels, et al.
Veröffentlicht: (2024)
FuzzAug: Data Augmentation by Coverage-guided Fuzzing for Neural Test Generation
von: He, Yifeng, et al.
Veröffentlicht: (2024)
von: He, Yifeng, et al.
Veröffentlicht: (2024)
ToolPRMBench: Evaluating and Advancing Process Reward Models for Tool-using Agents
von: Li, Dawei, et al.
Veröffentlicht: (2026)
von: Li, Dawei, et al.
Veröffentlicht: (2026)
Z-Space: A Multi-Agent Tool Orchestration Framework for Enterprise-Grade LLM Automation
von: He, Qingsong, et al.
Veröffentlicht: (2025)
von: He, Qingsong, et al.
Veröffentlicht: (2025)
Deploy-Master: Automating the Deployment of 50,000+ Agent-Ready Scientific Tools in One Day
von: Wang, Yi, et al.
Veröffentlicht: (2026)
von: Wang, Yi, et al.
Veröffentlicht: (2026)
Automating Android Build Repair: Bridging the Reasoning-Execution Gap in LLM Agents with Domain-Specific Tools
von: Son, Ha Min, et al.
Veröffentlicht: (2025)
von: Son, Ha Min, et al.
Veröffentlicht: (2025)
Evaluating AGENTS.md: Are Repository-Level Context Files Helpful for Coding Agents?
von: Gloaguen, Thibaud, et al.
Veröffentlicht: (2026)
von: Gloaguen, Thibaud, et al.
Veröffentlicht: (2026)
Schema First Tool APIs for LLM Agents: A Controlled Study of Tool Misuse, Recovery, and Budgeted Performance
von: Sigdel, Akshey, et al.
Veröffentlicht: (2026)
von: Sigdel, Akshey, et al.
Veröffentlicht: (2026)
Automated Creation and Enrichment Framework for Improved Invocation of Enterprise APIs as Tools
von: Agarwal, Prerna, et al.
Veröffentlicht: (2025)
von: Agarwal, Prerna, et al.
Veröffentlicht: (2025)
MCP-Zero: Active Tool Discovery for Autonomous LLM Agents
von: Fei, Xiang, et al.
Veröffentlicht: (2025)
von: Fei, Xiang, et al.
Veröffentlicht: (2025)
ASA: Training-Free Representation Engineering for Tool-Calling Agents
von: Wang, Youjin, et al.
Veröffentlicht: (2026)
von: Wang, Youjin, et al.
Veröffentlicht: (2026)
Squeez: Task-Conditioned Tool-Output Pruning for Coding Agents
von: Kovács, Ádám
Veröffentlicht: (2026)
von: Kovács, Ádám
Veröffentlicht: (2026)
Investigating Tool-Memory Conflicts in Tool-Augmented LLMs
von: Cheng, Jiali, et al.
Veröffentlicht: (2026)
von: Cheng, Jiali, et al.
Veröffentlicht: (2026)
A Tool for Generating Exceptional Behavior Tests With Large Language Models
von: Zhong, Linghan, et al.
Veröffentlicht: (2025)
von: Zhong, Linghan, et al.
Veröffentlicht: (2025)
Towards Reliable LLM-Driven Fuzz Testing: Vision and Road Ahead
von: Cheng, Yiran, et al.
Veröffentlicht: (2025)
von: Cheng, Yiran, et al.
Veröffentlicht: (2025)
ParaTool: Shifting Tool Representations from Context to Parameters
von: Yu, Zekai, et al.
Veröffentlicht: (2026)
von: Yu, Zekai, et al.
Veröffentlicht: (2026)
SynthTools: A Framework for Scaling Synthetic Tools for Agent Development
von: Castellani, Tommaso, et al.
Veröffentlicht: (2025)
von: Castellani, Tommaso, et al.
Veröffentlicht: (2025)
AutoSafeCoder: A Multi-Agent Framework for Securing LLM Code Generation through Static Analysis and Fuzz Testing
von: Nunez, Ana, et al.
Veröffentlicht: (2024)
von: Nunez, Ana, et al.
Veröffentlicht: (2024)
Solver-Aided Verification of Policy Compliance in Tool-Augmented LLM Agents
von: Winston, Cailin, et al.
Veröffentlicht: (2026)
von: Winston, Cailin, et al.
Veröffentlicht: (2026)
ToolScan: A Benchmark for Characterizing Errors in Tool-Use LLMs
von: Kokane, Shirley, et al.
Veröffentlicht: (2024)
von: Kokane, Shirley, et al.
Veröffentlicht: (2024)
JTPRO: A Joint Tool-Prompt Reflective Optimization Framework for Language Agents
von: Ghoshal, Sandip, et al.
Veröffentlicht: (2026)
von: Ghoshal, Sandip, et al.
Veröffentlicht: (2026)
Graph-Based Self-Healing Tool Routing for Cost-Efficient LLM Agents
von: Bholani, Neeraj
Veröffentlicht: (2026)
von: Bholani, Neeraj
Veröffentlicht: (2026)
Live API-Bench: 2500+ Live APIs for Testing Multi-Step Tool Calling
von: Elder, Benjamin, et al.
Veröffentlicht: (2025)
von: Elder, Benjamin, et al.
Veröffentlicht: (2025)
HFuzzer: Testing Large Language Models for Package Hallucinations via Phrase-based Fuzzing
von: Zhao, Yukai, et al.
Veröffentlicht: (2025)
von: Zhao, Yukai, et al.
Veröffentlicht: (2025)
LLMs are All You Need? Improving Fuzz Testing for MOJO with Large Language Models
von: Huang, Linghan, et al.
Veröffentlicht: (2025)
von: Huang, Linghan, et al.
Veröffentlicht: (2025)
EvolveTool-Bench: Evaluating the Quality of LLM-Generated Tool Libraries as Software Artifacts
von: Kaliyev, Alibek T., et al.
Veröffentlicht: (2026)
von: Kaliyev, Alibek T., et al.
Veröffentlicht: (2026)
ToolMisuseBench: An Offline Deterministic Benchmark for Tool Misuse and Recovery in Agentic Systems
von: Sigdel, Akshey, et al.
Veröffentlicht: (2026)
von: Sigdel, Akshey, et al.
Veröffentlicht: (2026)
Rethinking the Role of Entropy in Optimizing Tool-Use Behaviors for Large Language Model Agents
von: Li, Zeping, et al.
Veröffentlicht: (2026)
von: Li, Zeping, et al.
Veröffentlicht: (2026)
ComplexMCP: Evaluation of LLM Agents in Dynamic, Interdependent, and Large-Scale Tool Sandbox
von: Li, Yuanyang, et al.
Veröffentlicht: (2026)
von: Li, Yuanyang, et al.
Veröffentlicht: (2026)
MathViz-E: A Case-study in Domain-Specialized Tool-Using Agents
von: Bulusu, Arya, et al.
Veröffentlicht: (2024)
von: Bulusu, Arya, et al.
Veröffentlicht: (2024)
Open, Reliable, and Collective: A Community-Driven Framework for Tool-Using AI Agents
von: Dang, Hy, et al.
Veröffentlicht: (2026)
von: Dang, Hy, et al.
Veröffentlicht: (2026)
Unit Test Generation using Generative AI : A Comparative Performance Analysis of Autogeneration Tools
von: Bhatia, Shreya, et al.
Veröffentlicht: (2023)
von: Bhatia, Shreya, et al.
Veröffentlicht: (2023)
The Tool-Overuse Illusion: Why Does LLM Prefer External Tools over Internal Knowledge?
von: Zeng, Yirong, et al.
Veröffentlicht: (2026)
von: Zeng, Yirong, et al.
Veröffentlicht: (2026)
An Executable Benchmarking Suite for Tool-Using Agents
von: Zhong, Zhiqing, et al.
Veröffentlicht: (2026)
von: Zhong, Zhiqing, et al.
Veröffentlicht: (2026)
Comparison of Static Application Security Testing Tools and Large Language Models for Repo-level Vulnerability Detection
von: Zhou, Xin, et al.
Veröffentlicht: (2024)
von: Zhou, Xin, et al.
Veröffentlicht: (2024)
Semantic Tool Discovery for Large Language Models: A Vector-Based Approach to MCP Tool Selection
von: Mudunuri, Sarat, et al.
Veröffentlicht: (2026)
von: Mudunuri, Sarat, et al.
Veröffentlicht: (2026)
Breaking the Illusion of Identity in LLM Tooling
von: Miller, Marek
Veröffentlicht: (2026)
von: Miller, Marek
Veröffentlicht: (2026)
Ähnliche Einträge
-
Automated Benchmark Generation for Repository-Level Coding Tasks
von: Vergopoulos, Konstantinos, et al.
Veröffentlicht: (2025) -
AutoRestTest: A Tool for Automated REST API Testing Using LLMs and MARL
von: Stennett, Tyler, et al.
Veröffentlicht: (2025) -
MIMIC-Py: An Extensible Tool for Personality-Driven Automated Game Testing with Large Language Models
von: Chen, Yifei, et al.
Veröffentlicht: (2026) -
SWT-Bench: Testing and Validating Real-World Bug-Fixes with Code Agents
von: Mündler, Niels, et al.
Veröffentlicht: (2024) -
FuzzAug: Data Augmentation by Coverage-guided Fuzzing for Neural Test Generation
von: He, Yifeng, et al.
Veröffentlicht: (2024)