Evaluating LLMs on Sequential API Call Through Automated Test Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Yuheng, Song, Jiayang, Song, Da, Ji, Zhenlan, Wang, Wenhan, Wang, Shuai, Ma, Lei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TRUSTVIS: A Multi-Dimensional Trustworthiness Evaluation Framework for Large Language Models
by: Sun, Ruoyu, et al.
Published: (2025)
by: Sun, Ruoyu, et al.
Published: (2025)
Online Safety Analysis for LLMs: a Benchmark, an Assessment, and a Path Forward
by: Xie, Xuan, et al.
Published: (2024)
by: Xie, Xuan, et al.
Published: (2024)
Digging Into the Internal: Causality-Based Analysis of LLM Function Calling
by: Ji, Zhenlan, et al.
Published: (2025)
by: Ji, Zhenlan, et al.
Published: (2025)
LeCov: Multi-level Testing Criteria for Large Language Models
by: Xie, Xuan, et al.
Published: (2024)
by: Xie, Xuan, et al.
Published: (2024)
Look Before You Leap: An Exploratory Study of Uncertainty Measurement for Large Language Models
by: Huang, Yuheng, et al.
Published: (2023)
by: Huang, Yuheng, et al.
Published: (2023)
LUNA: A Model-Based Universal Analysis Framework for Large Language Models
by: Song, Da, et al.
Published: (2023)
by: Song, Da, et al.
Published: (2023)
AcTracer: Active Testing of Large Language Model via Multi-Stage Sampling
by: Huang, Yuheng, et al.
Published: (2024)
by: Huang, Yuheng, et al.
Published: (2024)
Evaluating Retrieval-Augmented Generation Variants for Natural Language-Based SQL and API Call Generation
by: Marketsmüller, Michael, et al.
Published: (2026)
by: Marketsmüller, Michael, et al.
Published: (2026)
CAM: A Causality-based Analysis Framework for Multi-Agent Code Generation Systems
by: Lyu, Zongyi, et al.
Published: (2026)
by: Lyu, Zongyi, et al.
Published: (2026)
Reverse Chain: A Generic-Rule for LLMs to Master Multi-API Planning
by: Zhang, Yinger, et al.
Published: (2023)
by: Zhang, Yinger, et al.
Published: (2023)
Risk Assessment Framework for Code LLMs via Leveraging Internal States
by: Huang, Yuheng, et al.
Published: (2025)
by: Huang, Yuheng, et al.
Published: (2025)
VLATest: Testing and Evaluating Vision-Language-Action Models for Robotic Manipulation
by: Wang, Zhijie, et al.
Published: (2024)
by: Wang, Zhijie, et al.
Published: (2024)
TESTEVAL: Benchmarking Large Language Models for Test Case Generation
by: Wang, Wenhan, et al.
Published: (2024)
by: Wang, Wenhan, et al.
Published: (2024)
Understanding the Fundamental Design Decisions of Retrieval-Augmented Generation Systems
by: Zhao, Shengming, et al.
Published: (2024)
by: Zhao, Shengming, et al.
Published: (2024)
Compositional API Recommendation for Library-Oriented Code Generation
by: Ma, Zexiong, et al.
Published: (2024)
by: Ma, Zexiong, et al.
Published: (2024)
Evaluating Implicit Regulatory Compliance in LLM Tool Invocation via Logic-Guided Synthesis
by: Song, Da, et al.
Published: (2026)
by: Song, Da, et al.
Published: (2026)
ToolFactory: Automating Tool Generation by Leveraging LLM to Understand REST API Documentations
by: Ni, Xinyi, et al.
Published: (2025)
by: Ni, Xinyi, et al.
Published: (2025)
AutoRestTest: A Tool for Automated REST API Testing Using LLMs and MARL
by: Stennett, Tyler, et al.
Published: (2025)
by: Stennett, Tyler, et al.
Published: (2025)
Fine-grained Testing for Autonomous Driving Software: a Study on Autoware with LLM-driven Unit Testing
by: Wang, Wenhan, et al.
Published: (2025)
by: Wang, Wenhan, et al.
Published: (2025)
Code-Vision: Evaluating Multimodal LLMs Logic Understanding and Code Generation Capabilities
by: Wang, Hanbin, et al.
Published: (2025)
by: Wang, Hanbin, et al.
Published: (2025)
Assessing Evaluation Metrics for Neural Test Oracle Generation
by: Shin, Jiho, et al.
Published: (2023)
by: Shin, Jiho, et al.
Published: (2023)
Understanding and Bridging the Planner-Coder Gap: A Systematic Study on the Robustness of Multi-Agent Systems for Code Generation
by: Lyu, Zongyi, et al.
Published: (2025)
by: Lyu, Zongyi, et al.
Published: (2025)
The Prompt Alchemist: Automated LLM-Tailored Prompt Optimization for Test Case Generation
by: Gao, Shuzheng, et al.
Published: (2025)
by: Gao, Shuzheng, et al.
Published: (2025)
Live API-Bench: 2500+ Live APIs for Testing Multi-Step Tool Calling
by: Elder, Benjamin, et al.
Published: (2025)
by: Elder, Benjamin, et al.
Published: (2025)
APIGen: Automated Pipeline for Generating Verifiable and Diverse Function-Calling Datasets
by: Liu, Zuxin, et al.
Published: (2024)
by: Liu, Zuxin, et al.
Published: (2024)
Can LLMs Generate Reliable Test Case Generators? A Study on Competition-Level Programming Problems
by: Cao, Yuhan, et al.
Published: (2025)
by: Cao, Yuhan, et al.
Published: (2025)
Automating a Complete Software Test Process Using LLMs: An Automotive Case Study
by: Wang, Shuai, et al.
Published: (2025)
by: Wang, Shuai, et al.
Published: (2025)
Towards Automated Smart Contract Generation: Evaluation, Benchmarking, and Retrieval-Augmented Repair
by: Chen, Zaoyu, et al.
Published: (2025)
by: Chen, Zaoyu, et al.
Published: (2025)
SoAy: A Solution-based LLM API-using Methodology for Academic Information Seeking
by: Wang, Yuanchun, et al.
Published: (2024)
by: Wang, Yuanchun, et al.
Published: (2024)
Python Symbolic Execution with LLM-powered Code Generation
by: Wang, Wenhan, et al.
Published: (2024)
by: Wang, Wenhan, et al.
Published: (2024)
Revolutionizing API Documentation through Summarization
by: Naghshzan, AmirHossein, et al.
Published: (2024)
by: Naghshzan, AmirHossein, et al.
Published: (2024)
MORTAR: A Model-based Runtime Action Repair Framework for AI-enabled Cyber-Physical Systems
by: Wang, Renzhi, et al.
Published: (2024)
by: Wang, Renzhi, et al.
Published: (2024)
Characterizing and Evaluating the Reliability of LLMs against Jailbreak Attacks
by: Chen, Kexin, et al.
Published: (2024)
by: Chen, Kexin, et al.
Published: (2024)
ACECODER: Acing Coder RL via Automated Test-Case Synthesis
by: Zeng, Huaye, et al.
Published: (2025)
by: Zeng, Huaye, et al.
Published: (2025)
Combining TSL and LLM to Automate REST API Testing: A Comparative Study
by: Barradas, Thiago, et al.
Published: (2025)
by: Barradas, Thiago, et al.
Published: (2025)
Testing and Evaluation of Large Language Models: Correctness, Non-Toxicity, and Fairness
by: Wang, Wenxuan
Published: (2024)
by: Wang, Wenxuan
Published: (2024)
CodeScope: An Execution-based Multilingual Multitask Multidimensional Benchmark for Evaluating LLMs on Code Understanding and Generation
by: Yan, Weixiang, et al.
Published: (2023)
by: Yan, Weixiang, et al.
Published: (2023)
Are LLMs Correctly Integrated into Software Systems?
by: Shao, Yuchen, et al.
Published: (2024)
by: Shao, Yuchen, et al.
Published: (2024)
Towards Understanding the Characteristics of Code Generation Errors Made by Large Language Models
by: Wang, Zhijie, et al.
Published: (2024)
by: Wang, Zhijie, et al.
Published: (2024)
Enhancing Large Language Models in Coding Through Multi-Perspective Self-Consistency
by: Huang, Baizhou, et al.
Published: (2023)
by: Huang, Baizhou, et al.
Published: (2023)
Similar Items
-
TRUSTVIS: A Multi-Dimensional Trustworthiness Evaluation Framework for Large Language Models
by: Sun, Ruoyu, et al.
Published: (2025) -
Online Safety Analysis for LLMs: a Benchmark, an Assessment, and a Path Forward
by: Xie, Xuan, et al.
Published: (2024) -
Digging Into the Internal: Causality-Based Analysis of LLM Function Calling
by: Ji, Zhenlan, et al.
Published: (2025) -
LeCov: Multi-level Testing Criteria for Large Language Models
by: Xie, Xuan, et al.
Published: (2024) -
Look Before You Leap: An Exploratory Study of Uncertainty Measurement for Large Language Models
by: Huang, Yuheng, et al.
Published: (2023)