StableToolBench-MirrorAPI: Modeling Tool Environments as Mirrors of 7,000+ Real-World APIs
Fuente:
arXiv
Salvato in:
| Autori principali: | Guo, Zhicheng, Cheng, Sijie, Niu, Yuchen, Wang, Hao, Zhou, Sicheng, Huang, Wenbing, Liu, Yang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
StableToolBench: Towards Stable Large-Scale Benchmarking on Tool Learning of Large Language Models
di: Guo, Zhicheng, et al.
Pubblicazione: (2024)
di: Guo, Zhicheng, et al.
Pubblicazione: (2024)
Live API-Bench: 2500+ Live APIs for Testing Multi-Step Tool Calling
di: Elder, Benjamin, et al.
Pubblicazione: (2025)
di: Elder, Benjamin, et al.
Pubblicazione: (2025)
MirrorBench: Evaluating Self-centric Intelligence in MLLMs by Introducing a Mirror
di: Guo, Shengyu, et al.
Pubblicazione: (2026)
di: Guo, Shengyu, et al.
Pubblicazione: (2026)
HarnessAPI: A Skill-First Framework for Unified Streaming APIs and MCP Tools
di: Jose, Edwin
Pubblicazione: (2026)
di: Jose, Edwin
Pubblicazione: (2026)
Beyond Perfect APIs: A Comprehensive Evaluation of LLM Agents Under Real-World API Complexity
di: Kim, Doyoung, et al.
Pubblicazione: (2026)
di: Kim, Doyoung, et al.
Pubblicazione: (2026)
ToolForge: A Data Synthesis Pipeline for Multi-Hop Search without Real-World APIs
di: Chen, Hao, et al.
Pubblicazione: (2025)
di: Chen, Hao, et al.
Pubblicazione: (2025)
Code2API: A Tool for Generating Reusable APIs from Stack Overflow Code Snippets
di: Mai, Yubo, et al.
Pubblicazione: (2025)
di: Mai, Yubo, et al.
Pubblicazione: (2025)
FinToolBench: Evaluating LLM Agents for Real-World Financial Tool Use
di: Lu, Jiaxuan, et al.
Pubblicazione: (2026)
di: Lu, Jiaxuan, et al.
Pubblicazione: (2026)
MirrorFuzz: Leveraging LLM and Shared Bugs for Deep Learning Framework APIs Fuzzing
di: Ou, Shiwen, et al.
Pubblicazione: (2025)
di: Ou, Shiwen, et al.
Pubblicazione: (2025)
API Security: Protecting APIs With Keycloak
di: Danso Solomon Danquah, et al.
Pubblicazione: (2023)
di: Danso Solomon Danquah, et al.
Pubblicazione: (2023)
Mirror of Nature, Mirror of Self
di: Shevchenko, Dimitry
Pubblicazione: (2024)
di: Shevchenko, Dimitry
Pubblicazione: (2024)
Foundational VeriFast: Pragmatic Certification of Verification Tool Results through Hinted Mirroring
di: Jacobs, Bart
Pubblicazione: (2026)
di: Jacobs, Bart
Pubblicazione: (2026)
Algorithmic Mirror: Designing an Interactive Tool to Promote Self-Reflection for YouTube Recommendations
di: Kondo, Yui, et al.
Pubblicazione: (2025)
di: Kondo, Yui, et al.
Pubblicazione: (2025)
Stable Phase Retrieval with Mirror Descent
di: Godeme, Jean-Jacques, et al.
Pubblicazione: (2024)
di: Godeme, Jean-Jacques, et al.
Pubblicazione: (2024)
MirrorLimb: Implementing hand pose acquisition and robot teleoperation based on RealMirror
di: Tai, Cong, et al.
Pubblicazione: (2025)
di: Tai, Cong, et al.
Pubblicazione: (2025)
Firefly: Illuminating Large-Scale Verified Tool-Call Data Generation from Real APIs
di: Lu, Yuxuan, et al.
Pubblicazione: (2026)
di: Lu, Yuxuan, et al.
Pubblicazione: (2026)
MirrorGaussian: Reflecting 3D Gaussians for Reconstructing Mirror Reflections
di: Liu, Jiayue, et al.
Pubblicazione: (2024)
di: Liu, Jiayue, et al.
Pubblicazione: (2024)
Notes on GLSMs for Supermanifolds and Their Mirrors
di: Zou, Hao
Pubblicazione: (2025)
di: Zou, Hao
Pubblicazione: (2025)
Mirror, Mirror of the Flow: How Does Regularization Shape Implicit Bias?
di: Jacobs, Tom, et al.
Pubblicazione: (2025)
di: Jacobs, Tom, et al.
Pubblicazione: (2025)
MCP-AgentBench: Evaluating Real-World Language Agent Performance with MCP-Mediated Tools
di: Guo, Zikang, et al.
Pubblicazione: (2025)
di: Guo, Zikang, et al.
Pubblicazione: (2025)
Mirror
di: Francisca Labbé
Pubblicazione: (2003)
di: Francisca Labbé
Pubblicazione: (2003)
Walls Are Mirrors: Messages Delivered in Physical Environments
di: Heeyoung Han, et al.
Pubblicazione: (2025)
di: Heeyoung Han, et al.
Pubblicazione: (2025)
3d Mirror Symmetry is Mirror Symmetry
di: Chan, Ki Fung, et al.
Pubblicazione: (2024)
di: Chan, Ki Fung, et al.
Pubblicazione: (2024)
Embodied Representation Alignment with Mirror Neurons
di: Zhu, Wentao, et al.
Pubblicazione: (2025)
di: Zhu, Wentao, et al.
Pubblicazione: (2025)
Matter-Dark Matter Coincidence and Mirror World
di: Mohapatra, Rabindra N., et al.
Pubblicazione: (2025)
di: Mohapatra, Rabindra N., et al.
Pubblicazione: (2025)
A Framework for Testing and Adapting REST APIs as LLM Tools
di: Bandlamudi, Jayachandu, et al.
Pubblicazione: (2025)
di: Bandlamudi, Jayachandu, et al.
Pubblicazione: (2025)
$G_2$ Mirrors from Calabi-Yau Mirrors
di: Braun, Andreas P., et al.
Pubblicazione: (2023)
di: Braun, Andreas P., et al.
Pubblicazione: (2023)
Mirror, Mirror on the Wall: Reflecting the Best in Popular Culture.
di: Taylor, Rhonda Harris, et al.
Pubblicazione: (1996)
di: Taylor, Rhonda Harris, et al.
Pubblicazione: (1996)
WorldMirror: Universal 3D World Reconstruction with Any-Prior Prompting
di: Liu, Yifan, et al.
Pubblicazione: (2025)
di: Liu, Yifan, et al.
Pubblicazione: (2025)
Design and Test of Small Mirror Supports for Harsh Environments
di: Huie, Ruby, et al.
Pubblicazione: (2024)
di: Huie, Ruby, et al.
Pubblicazione: (2024)
Toward Real-Time Mirrors Intelligence: System-Level Latency and Computation Evaluation in Internet of Mirrors (IoM)
di: Fatima, Haneen, et al.
Pubblicazione: (2026)
di: Fatima, Haneen, et al.
Pubblicazione: (2026)
Mirror-world
di: Cho, Daniel
Pubblicazione: (2025)
di: Cho, Daniel
Pubblicazione: (2025)
Documenti analoghi
-
StableToolBench: Towards Stable Large-Scale Benchmarking on Tool Learning of Large Language Models
di: Guo, Zhicheng, et al.
Pubblicazione: (2024) -
Live API-Bench: 2500+ Live APIs for Testing Multi-Step Tool Calling
di: Elder, Benjamin, et al.
Pubblicazione: (2025) -
MirrorBench: Evaluating Self-centric Intelligence in MLLMs by Introducing a Mirror
di: Guo, Shengyu, et al.
Pubblicazione: (2026) -
HarnessAPI: A Skill-First Framework for Unified Streaming APIs and MCP Tools
di: Jose, Edwin
Pubblicazione: (2026) -
Beyond Perfect APIs: A Comprehensive Evaluation of LLM Agents Under Real-World API Complexity
di: Kim, Doyoung, et al.
Pubblicazione: (2026)