Beyond Perfect APIs: A Comprehensive Evaluation of LLM Agents Under Real-World API Complexity
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, Doyoung, Ren, Zhiwei, Hao, Jie, Sun, Zhongkai, Wang, Lichao, Ma, Xiyao, Ye, Zack, Han, Xu, Yin, Jun, Ji, Heng, Shen, Wei, Fan, Xing, Yao, Benjamin, Guo, Chenlei |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
API Security: Protecting APIs With Keycloak
by: Danso Solomon Danquah, et al.
Published: (2023)
by: Danso Solomon Danquah, et al.
Published: (2023)
Analyzing and Internalizing Complex Policy Documents for LLM Agents
by: Liu, Jiateng, et al.
Published: (2025)
by: Liu, Jiateng, et al.
Published: (2025)
StableToolBench-MirrorAPI: Modeling Tool Environments as Mirrors of 7,000+ Real-World APIs
by: Guo, Zhicheng, et al.
Published: (2025)
by: Guo, Zhicheng, et al.
Published: (2025)
WorldAPIs: The World Is Worth How Many APIs? A Thought Experiment
by: Ou, Jiefu, et al.
Published: (2024)
by: Ou, Jiefu, et al.
Published: (2024)
Beyond SELECT: A Comprehensive Taxonomy-Guided Benchmark for Real-World Text-to-SQL Translation
by: Wang, Hao, et al.
Published: (2025)
by: Wang, Hao, et al.
Published: (2025)
Live API-Bench: 2500+ Live APIs for Testing Multi-Step Tool Calling
by: Elder, Benjamin, et al.
Published: (2025)
by: Elder, Benjamin, et al.
Published: (2025)
Beyond Indistinguishability: Measuring Extraction Risk in LLM APIs
by: Liu, Ruixuan, et al.
Published: (2026)
by: Liu, Ruixuan, et al.
Published: (2026)
Code2API: A Tool for Generating Reusable APIs from Stack Overflow Code Snippets
by: Mai, Yubo, et al.
Published: (2025)
by: Mai, Yubo, et al.
Published: (2025)
Pricing4APIs: A Rigorous Model for RESTful API Pricings
by: Fresno-Aranda, Rafael, et al.
Published: (2023)
by: Fresno-Aranda, Rafael, et al.
Published: (2023)
Semantic API Alignment: Linking High-level User Goals to APIs
by: Feldt, Robert, et al.
Published: (2024)
by: Feldt, Robert, et al.
Published: (2024)
PA3: Policy-Aware Agent Alignment through Chain-of-Thought
by: Dipta, Shubhashis Roy, et al.
Published: (2026)
by: Dipta, Shubhashis Roy, et al.
Published: (2026)
Do RESTful API Design Rules Have an Impact on the Understandability of Web APIs? A Web-Based Experiment with API Descriptions
by: Bogner, Justus, et al.
Published: (2023)
by: Bogner, Justus, et al.
Published: (2023)
Beyond Browsing: API-Based Web Agents
by: Song, Yueqi, et al.
Published: (2024)
by: Song, Yueqi, et al.
Published: (2024)
HarnessAPI: A Skill-First Framework for Unified Streaming APIs and MCP Tools
by: Jose, Edwin
Published: (2026)
by: Jose, Edwin
Published: (2026)
AI Agentic workflows and Enterprise APIs: Adapting API architectures for the age of AI agents
by: Tupe, Vaibhav, et al.
Published: (2025)
by: Tupe, Vaibhav, et al.
Published: (2025)
Multimodal Policy Internalization for Conversational Agents
by: Wang, Zhenhailong, et al.
Published: (2025)
by: Wang, Zhenhailong, et al.
Published: (2025)
ToolForge: A Data Synthesis Pipeline for Multi-Hop Search without Real-World APIs
by: Chen, Hao, et al.
Published: (2025)
by: Chen, Hao, et al.
Published: (2025)
Beyond Text-to-SQL: An Agentic LLM System for Governed Enterprise Analytics APIs
by: Singh, Gundeep, et al.
Published: (2026)
by: Singh, Gundeep, et al.
Published: (2026)
Singularities of the nested Hilbert scheme of points of length 3, 4
by: Choi, Doyoung
Published: (2025)
by: Choi, Doyoung
Published: (2025)
Degree of the 3-secant variety
by: Choi, Doyoung
Published: (2023)
by: Choi, Doyoung
Published: (2023)
MEND: Meta dEmonstratioN Distillation for Efficient and Effective In-Context Learning
by: Li, Yichuan, et al.
Published: (2024)
by: Li, Yichuan, et al.
Published: (2024)
Towards Understanding Android APIs: Official Lists, Vendor Customizations, and Real-World Usage
by: Wang, Sinan, et al.
Published: (2026)
by: Wang, Sinan, et al.
Published: (2026)
Log Probability Tracking of LLM APIs
by: Chauvin, Timothée, et al.
Published: (2025)
by: Chauvin, Timothée, et al.
Published: (2025)
Test Amplification for REST APIs via Single and Multi-Agent LLM Systems
by: Nooyens, Robbe, et al.
Published: (2025)
by: Nooyens, Robbe, et al.
Published: (2025)
RAS-Eval: A Comprehensive Benchmark for Security Evaluation of LLM Agents in Real-World Environments
by: Fu, Yuchuan, et al.
Published: (2025)
by: Fu, Yuchuan, et al.
Published: (2025)
LiveClawBench: Benchmarking LLM Agents on Complex, Real-World Assistant Tasks
by: Long, Xiang, et al.
Published: (2026)
by: Long, Xiang, et al.
Published: (2026)
From Summary to Action: Enhancing Large Language Models for Complex Tasks with Open World APIs
by: Liu, Yulong, et al.
Published: (2024)
by: Liu, Yulong, et al.
Published: (2024)
Token-Efficient Change Detection in LLM APIs
by: Chauvin, Timothée, et al.
Published: (2026)
by: Chauvin, Timothée, et al.
Published: (2026)
Toward Real-World Table Agents: Capabilities, Workflows, and Design Principles for LLM-based Table Intelligence
by: Tian, Jiaming, et al.
Published: (2025)
by: Tian, Jiaming, et al.
Published: (2025)
The Perfection Paradox: From Architect to Curator in AI-Assisted API Design
by: Ahmad, Mak, et al.
Published: (2026)
by: Ahmad, Mak, et al.
Published: (2026)
CoIn: Counting the Invisible Reasoning Tokens in Commercial Opaque LLM APIs
by: Sun, Guoheng, et al.
Published: (2025)
by: Sun, Guoheng, et al.
Published: (2025)
MolLingo: Molecule-Native Representations for LLM-Powered Scientific Agents
by: Nguyen, Thao, et al.
Published: (2026)
by: Nguyen, Thao, et al.
Published: (2026)
API-BLEND: A Comprehensive Corpora for Training and Benchmarking API LLMs
by: Basu, Kinjal, et al.
Published: (2024)
by: Basu, Kinjal, et al.
Published: (2024)
HWE-Bench: Benchmarking LLM Agents on Real-World Hardware Bug Repair Tasks
by: Cui, Fan, et al.
Published: (2026)
by: Cui, Fan, et al.
Published: (2026)
SafeMCP: Proactive Power Regulation for LLM Agent Defense via Environment-Grounded Look-Ahead Reasoning
by: Wang, Lichao, et al.
Published: (2026)
by: Wang, Lichao, et al.
Published: (2026)
CirrusBench: Evaluating LLM-based Agents Beyond Correctness in Real-World Cloud Service Environments
by: Yu, Yi, et al.
Published: (2026)
by: Yu, Yi, et al.
Published: (2026)
Horizon-LM: A RAM-Centric Architecture for LLM Training
by: Yuan, Zhengqing, et al.
Published: (2026)
by: Yuan, Zhengqing, et al.
Published: (2026)
Perfect Worlds
by: Fokkema, Douwe
Published: (2011)
by: Fokkema, Douwe
Published: (2011)
GUARDIAN: Safeguarding LLM Multi-Agent Collaborations with Temporal Graph Modeling
by: Zhou, Jialong, et al.
Published: (2025)
by: Zhou, Jialong, et al.
Published: (2025)
DataGovBench: Benchmarking LLM Agents for Real-World Data Governance Workflows
by: Liu, Zhou, et al.
Published: (2025)
by: Liu, Zhou, et al.
Published: (2025)
Similar Items
-
API Security: Protecting APIs With Keycloak
by: Danso Solomon Danquah, et al.
Published: (2023) -
Analyzing and Internalizing Complex Policy Documents for LLM Agents
by: Liu, Jiateng, et al.
Published: (2025) -
StableToolBench-MirrorAPI: Modeling Tool Environments as Mirrors of 7,000+ Real-World APIs
by: Guo, Zhicheng, et al.
Published: (2025) -
WorldAPIs: The World Is Worth How Many APIs? A Thought Experiment
by: Ou, Jiefu, et al.
Published: (2024) -
Beyond SELECT: A Comprehensive Taxonomy-Guided Benchmark for Real-World Text-to-SQL Translation
by: Wang, Hao, et al.
Published: (2025)