HammerBench: Fine-Grained Function-Calling Evaluation in Real Mobile Device Scenarios

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Jun, Zhou, Jiamu, Wen, Muning, Mo, Xiaoyun, Zhang, Haoyu, Lin, Qiqiang, Jin, Cheng, Wang, Xihuai, Zhang, Weinan, Peng, Qiuying
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916616670806016
author Wang, Jun
Zhou, Jiamu
Wen, Muning
Mo, Xiaoyun
Zhang, Haoyu
Lin, Qiqiang
Jin, Cheng
Wang, Xihuai
Zhang, Weinan
Peng, Qiuying
Wang, Jun
author_facet Wang, Jun
Zhou, Jiamu
Wen, Muning
Mo, Xiaoyun
Zhang, Haoyu
Lin, Qiqiang
Jin, Cheng
Wang, Xihuai
Zhang, Weinan
Peng, Qiuying
Wang, Jun
contents Evaluating the performance of LLMs in multi-turn human-agent interactions presents significant challenges, particularly due to the complexity and variability of user behavior. In this paper, we introduce HammerBench, a novel benchmark framework for assessing LLMs' function-calling capabilities in real-world, multi-turn dialogues. HammerBench simulates diverse mobile assistant use cases, incorporating imperfect instructions, dynamic question-answer trajectories, intent and argument shifts, and the indirect use of external information through pronouns. To construct this benchmark, we curate a comprehensive dataset derived from popular mobile app functionalities and anonymized user logs, complemented by a cost-effective data generation pipeline leveraging open-source models. HammerBench is further augmented with fine-grained interaction snapshots and metrics, enabling detailed evaluation of function-calling performance across individual conversational turns. We demonstrate the effectiveness of HammerBench by evaluating several leading LLMs and uncovering key performance trends. Our experiments reveal that different types of parameter name errors are a significant source of failure across different interaction scenarios, highlighting critical areas for further improvement in LLM robustness for mobile assistant applications.
format Preprint
id arxiv_https___arxiv_org_abs_2412_16516
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle HammerBench: Fine-Grained Function-Calling Evaluation in Real Mobile Device Scenarios
Wang, Jun
Zhou, Jiamu
Wen, Muning
Mo, Xiaoyun
Zhang, Haoyu
Lin, Qiqiang
Jin, Cheng
Wang, Xihuai
Zhang, Weinan
Peng, Qiuying
Wang, Jun
Computation and Language
Human-Computer Interaction
Evaluating the performance of LLMs in multi-turn human-agent interactions presents significant challenges, particularly due to the complexity and variability of user behavior. In this paper, we introduce HammerBench, a novel benchmark framework for assessing LLMs' function-calling capabilities in real-world, multi-turn dialogues. HammerBench simulates diverse mobile assistant use cases, incorporating imperfect instructions, dynamic question-answer trajectories, intent and argument shifts, and the indirect use of external information through pronouns. To construct this benchmark, we curate a comprehensive dataset derived from popular mobile app functionalities and anonymized user logs, complemented by a cost-effective data generation pipeline leveraging open-source models. HammerBench is further augmented with fine-grained interaction snapshots and metrics, enabling detailed evaluation of function-calling performance across individual conversational turns. We demonstrate the effectiveness of HammerBench by evaluating several leading LLMs and uncovering key performance trends. Our experiments reveal that different types of parameter name errors are a significant source of failure across different interaction scenarios, highlighting critical areas for further improvement in LLM robustness for mobile assistant applications.
title HammerBench: Fine-Grained Function-Calling Evaluation in Real Mobile Device Scenarios
topic Computation and Language
Human-Computer Interaction
url https://arxiv.org/abs/2412.16516