EvolveTool-Bench: Evaluating the Quality of LLM-Generated Tool Libraries as Software Artifacts
Fuente:
arXiv
Saved in:
| Main Authors: | Kaliyev, Alibek T., Maryanskyy, Artem |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Evaluating LLM-Based 0-to-1 Software Generation in End-to-End CLI Tool Scenarios
by: Hu, Ruida, et al.
Published: (2026)
by: Hu, Ruida, et al.
Published: (2026)
Automated Validation of LLM-based Evaluators for Software Engineering Artifacts
by: Fandina, Ora Nova, et al.
Published: (2025)
by: Fandina, Ora Nova, et al.
Published: (2025)
ToolMisuseBench: An Offline Deterministic Benchmark for Tool Misuse and Recovery in Agentic Systems
by: Sigdel, Akshey, et al.
Published: (2026)
by: Sigdel, Akshey, et al.
Published: (2026)
Measuring LLM Trust Allocation Across Conflicting Software Artifacts
by: Ulfat, Noshin, et al.
Published: (2026)
by: Ulfat, Noshin, et al.
Published: (2026)
AI-Driven Tools in Modern Software Quality Assurance: An Assessment of Benefits, Challenges, and Future Directions
by: Pysmennyi, Ihor, et al.
Published: (2025)
by: Pysmennyi, Ihor, et al.
Published: (2025)
User Centric Evaluation of Code Generation Tools
by: Miah, Tanha, et al.
Published: (2024)
by: Miah, Tanha, et al.
Published: (2024)
Benchmarking AI Models in Software Engineering: A Review, Search Tool, and Unified Approach for Elevating Benchmark Quality
by: Koohestani, Roham, et al.
Published: (2025)
by: Koohestani, Roham, et al.
Published: (2025)
Breaking the Illusion of Identity in LLM Tooling
by: Miller, Marek
Published: (2026)
by: Miller, Marek
Published: (2026)
Integrating Various Software Artifacts for Better LLM-based Bug Localization and Program Repair
by: Feng, Qiong, et al.
Published: (2024)
by: Feng, Qiong, et al.
Published: (2024)
ToolPRMBench: Evaluating and Advancing Process Reward Models for Tool-using Agents
by: Li, Dawei, et al.
Published: (2026)
by: Li, Dawei, et al.
Published: (2026)
The Tool-Overuse Illusion: Why Does LLM Prefer External Tools over Internal Knowledge?
by: Zeng, Yirong, et al.
Published: (2026)
by: Zeng, Yirong, et al.
Published: (2026)
ComplexMCP: Evaluation of LLM Agents in Dynamic, Interdependent, and Large-Scale Tool Sandbox
by: Li, Yuanyang, et al.
Published: (2026)
by: Li, Yuanyang, et al.
Published: (2026)
Self-Evolving Software Agents
by: Robol, Marco, et al.
Published: (2026)
by: Robol, Marco, et al.
Published: (2026)
Live API-Bench: 2500+ Live APIs for Testing Multi-Step Tool Calling
by: Elder, Benjamin, et al.
Published: (2025)
by: Elder, Benjamin, et al.
Published: (2025)
Schema First Tool APIs for LLM Agents: A Controlled Study of Tool Misuse, Recovery, and Budgeted Performance
by: Sigdel, Akshey, et al.
Published: (2026)
by: Sigdel, Akshey, et al.
Published: (2026)
Evaluating LLM-Based Test Generation Under Software Evolution
by: Haroon, Sabaat, et al.
Published: (2026)
by: Haroon, Sabaat, et al.
Published: (2026)
ToolFuzz -- Automated Agent Tool Testing
by: Milev, Ivan, et al.
Published: (2025)
by: Milev, Ivan, et al.
Published: (2025)
RAG-MCP: Mitigating Prompt Bloat in LLM Tool Selection via Retrieval-Augmented Generation
by: Gan, Tiantian, et al.
Published: (2025)
by: Gan, Tiantian, et al.
Published: (2025)
EvoCodeBench: A Human-Performance Benchmark for Self-Evolving LLM-Driven Coding Systems
by: Zhang, Wentao, et al.
Published: (2026)
by: Zhang, Wentao, et al.
Published: (2026)
AISysRev -- LLM-based Tool for Title-abstract Screening
by: Huotala, Aleksi, et al.
Published: (2025)
by: Huotala, Aleksi, et al.
Published: (2025)
MCP-Zero: Active Tool Discovery for Autonomous LLM Agents
by: Fei, Xiang, et al.
Published: (2025)
by: Fei, Xiang, et al.
Published: (2025)
Open-Source AI-based SE Tools: Opportunities and Challenges of Collaborative Software Learning
by: Lin, Zhihao, et al.
Published: (2024)
by: Lin, Zhihao, et al.
Published: (2024)
Investigating Tool-Memory Conflicts in Tool-Augmented LLMs
by: Cheng, Jiali, et al.
Published: (2026)
by: Cheng, Jiali, et al.
Published: (2026)
ParaTool: Shifting Tool Representations from Context to Parameters
by: Yu, Zekai, et al.
Published: (2026)
by: Yu, Zekai, et al.
Published: (2026)
Solver-Aided Verification of Policy Compliance in Tool-Augmented LLM Agents
by: Winston, Cailin, et al.
Published: (2026)
by: Winston, Cailin, et al.
Published: (2026)
LoCoBench-Agent: An Interactive Benchmark for LLM Agents in Long-Context Software Engineering
by: Qiu, Jielin, et al.
Published: (2025)
by: Qiu, Jielin, et al.
Published: (2025)
ToolScan: A Benchmark for Characterizing Errors in Tool-Use LLMs
by: Kokane, Shirley, et al.
Published: (2024)
by: Kokane, Shirley, et al.
Published: (2024)
RoadmapBench: Evaluating Long-Horizon Agentic Software Development Across Version Upgrades
by: Xu, Xinbo, et al.
Published: (2026)
by: Xu, Xinbo, et al.
Published: (2026)
Copilot Evaluation Harness: Evaluating LLM-Guided Software Programming
by: Agarwal, Anisha, et al.
Published: (2024)
by: Agarwal, Anisha, et al.
Published: (2024)
Graph-Based Self-Healing Tool Routing for Cost-Efficient LLM Agents
by: Bholani, Neeraj
Published: (2026)
by: Bholani, Neeraj
Published: (2026)
AI-Driven Self-Evolving Software: A Promising Path Toward Software Automation
by: Cai, Liyi, et al.
Published: (2025)
by: Cai, Liyi, et al.
Published: (2025)
LogDx-CI: Benchmarking Log Reduction Tools for LLM Root-Cause Diagnosis
by: Qin, Bowen
Published: (2026)
by: Qin, Bowen
Published: (2026)
A Tool for Generating Exceptional Behavior Tests With Large Language Models
by: Zhong, Linghan, et al.
Published: (2025)
by: Zhong, Linghan, et al.
Published: (2025)
PyBench: Evaluating LLM Agent on various real-world coding tasks
by: Zhang, Yaolun, et al.
Published: (2024)
by: Zhang, Yaolun, et al.
Published: (2024)
Enhancing LLM-Based Coding Tools through Native Integration of IDE-Derived Static Context
by: Li, Yichen, et al.
Published: (2024)
by: Li, Yichen, et al.
Published: (2024)
Z-Space: A Multi-Agent Tool Orchestration Framework for Enterprise-Grade LLM Automation
by: He, Qingsong, et al.
Published: (2025)
by: He, Qingsong, et al.
Published: (2025)
Can LLM Generate Regression Tests for Software Commits?
by: Liu, Jing, et al.
Published: (2025)
by: Liu, Jing, et al.
Published: (2025)
RESTestBench: A Benchmark for Evaluating the Effectiveness of LLM-Generated REST API Test Cases from NL Requirements
by: Kogler, Leon, et al.
Published: (2026)
by: Kogler, Leon, et al.
Published: (2026)
ToolFactory: Automating Tool Generation by Leveraging LLM to Understand REST API Documentations
by: Ni, Xinyi, et al.
Published: (2025)
by: Ni, Xinyi, et al.
Published: (2025)
Semantic Tool Discovery for Large Language Models: A Vector-Based Approach to MCP Tool Selection
by: Mudunuri, Sarat, et al.
Published: (2026)
by: Mudunuri, Sarat, et al.
Published: (2026)
Similar Items
-
Evaluating LLM-Based 0-to-1 Software Generation in End-to-End CLI Tool Scenarios
by: Hu, Ruida, et al.
Published: (2026) -
Automated Validation of LLM-based Evaluators for Software Engineering Artifacts
by: Fandina, Ora Nova, et al.
Published: (2025) -
ToolMisuseBench: An Offline Deterministic Benchmark for Tool Misuse and Recovery in Agentic Systems
by: Sigdel, Akshey, et al.
Published: (2026) -
Measuring LLM Trust Allocation Across Conflicting Software Artifacts
by: Ulfat, Noshin, et al.
Published: (2026) -
AI-Driven Tools in Modern Software Quality Assurance: An Assessment of Benefits, Challenges, and Future Directions
by: Pysmennyi, Ihor, et al.
Published: (2025)