Live API-Bench: 2500+ Live APIs for Testing Multi-Step Tool Calling

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Elder, Benjamin, Murthi, Anupama, Kang, Jungkoo, Naik, Ankita Rajaram, Kate, Kiran, Basu, Kinjal, Contractor, Danish
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909999428534272
author Elder, Benjamin
Murthi, Anupama
Kang, Jungkoo
Naik, Ankita Rajaram
Kate, Kiran
Basu, Kinjal
Contractor, Danish
author_facet Elder, Benjamin
Murthi, Anupama
Kang, Jungkoo
Naik, Ankita Rajaram
Kate, Kiran
Basu, Kinjal
Contractor, Danish
contents Large language models (LLMs) increasingly rely on external tools and APIs to execute complex tasks specified in natural language. Evaluating such tool calling capabilities in realistic enterprise settings is challenging: APIs are often proprietary, heterogeneous, and difficult to share, limiting reproducible benchmarks. To address this, we introduce Live API Bench, a comprehensive benchmark constructed by transforming NL2SQL datasets into interactive API environments. Our pipeline converts SQL queries from BIRD SQL into executable API sequences across three formulations SLOT, SEL, and REST covering minimal general purpose operations, domain specific multi step tasks, and function oriented RESTful interactions, respectively. The benchmark spans 11 databases with over 2,500 invocable tools, paired with human authored queries, ground truth API sequences, and verified final answers. Live API Bench enables systematic evaluation of core challenges in tool use, including error handling, sequential reasoning, parameter generation, response parsing, and robustness across diverse domains. We evaluate 10 LLMs and 4 ReACT agents, observing low task completion rates (7 to 47pct), which improve modestly to 50pct under interactive agent settings, highlighting substantial scope for improving LLM tool calling performance. We release all code and data associated with this paper.
format Preprint
id arxiv_https___arxiv_org_abs_2506_11266
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Live API-Bench: 2500+ Live APIs for Testing Multi-Step Tool Calling
Elder, Benjamin
Murthi, Anupama
Kang, Jungkoo
Naik, Ankita Rajaram
Kate, Kiran
Basu, Kinjal
Contractor, Danish
Software Engineering
Artificial Intelligence
Large language models (LLMs) increasingly rely on external tools and APIs to execute complex tasks specified in natural language. Evaluating such tool calling capabilities in realistic enterprise settings is challenging: APIs are often proprietary, heterogeneous, and difficult to share, limiting reproducible benchmarks. To address this, we introduce Live API Bench, a comprehensive benchmark constructed by transforming NL2SQL datasets into interactive API environments. Our pipeline converts SQL queries from BIRD SQL into executable API sequences across three formulations SLOT, SEL, and REST covering minimal general purpose operations, domain specific multi step tasks, and function oriented RESTful interactions, respectively. The benchmark spans 11 databases with over 2,500 invocable tools, paired with human authored queries, ground truth API sequences, and verified final answers. Live API Bench enables systematic evaluation of core challenges in tool use, including error handling, sequential reasoning, parameter generation, response parsing, and robustness across diverse domains. We evaluate 10 LLMs and 4 ReACT agents, observing low task completion rates (7 to 47pct), which improve modestly to 50pct under interactive agent settings, highlighting substantial scope for improving LLM tool calling performance. We release all code and data associated with this paper.
title Live API-Bench: 2500+ Live APIs for Testing Multi-Step Tool Calling
topic Software Engineering
Artificial Intelligence
url https://arxiv.org/abs/2506.11266