CONFETTI: Conversational Function-Calling Evaluation Through Turn-Level Interactions

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Alkhouli, Tamer, Margatina, Katerina, Gung, James, Shu, Raphael, Zaghi, Claudia, Sunkara, Monica, Zhang, Yi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912409588858880
author Alkhouli, Tamer
Margatina, Katerina
Gung, James
Shu, Raphael
Zaghi, Claudia
Sunkara, Monica
Zhang, Yi
author_facet Alkhouli, Tamer
Margatina, Katerina
Gung, James
Shu, Raphael
Zaghi, Claudia
Sunkara, Monica
Zhang, Yi
contents We introduce Conversational Function-Calling Evaluation Through Turn-Level Interactions (CONFETTI), a conversational benchmark1 designed to evaluate the function-calling capabilities and response quality of large language models (LLMs). Current benchmarks lack comprehensive assessment of LLMs in complex conversational scenarios. CONFETTI addresses this gap through 109 human-simulated conversations, comprising 313 user turns and covering 86 APIs. These conversations explicitly target various conversational complexities, such as follow-ups, goal correction and switching, ambiguous and implicit goals. We perform off-policy turn-level evaluation using this benchmark targeting function-calling. Our benchmark also incorporates dialog act annotations to assess agent responses. We evaluate a series of state-of-the-art LLMs and analyze their performance with respect to the number of available APIs, conversation lengths, and chained function calling. Our results reveal that while some models are able to handle long conversations, and leverage more than 20+ APIs successfully, other models struggle with longer context or when increasing the number of APIs. We also report that the performance on chained function-calls is severely limited across the models. Overall, the top performing models on CONFETTI are Nova Pro (40.01%), Claude Sonnet v3.5 (35.46%) and Llama 3.1 405B (33.19%) followed by command-r-plus (31.18%) and Mistral-Large-2407 (30.07%).
format Preprint
id arxiv_https___arxiv_org_abs_2506_01859
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CONFETTI: Conversational Function-Calling Evaluation Through Turn-Level Interactions
Alkhouli, Tamer
Margatina, Katerina
Gung, James
Shu, Raphael
Zaghi, Claudia
Sunkara, Monica
Zhang, Yi
Computation and Language
We introduce Conversational Function-Calling Evaluation Through Turn-Level Interactions (CONFETTI), a conversational benchmark1 designed to evaluate the function-calling capabilities and response quality of large language models (LLMs). Current benchmarks lack comprehensive assessment of LLMs in complex conversational scenarios. CONFETTI addresses this gap through 109 human-simulated conversations, comprising 313 user turns and covering 86 APIs. These conversations explicitly target various conversational complexities, such as follow-ups, goal correction and switching, ambiguous and implicit goals. We perform off-policy turn-level evaluation using this benchmark targeting function-calling. Our benchmark also incorporates dialog act annotations to assess agent responses. We evaluate a series of state-of-the-art LLMs and analyze their performance with respect to the number of available APIs, conversation lengths, and chained function calling. Our results reveal that while some models are able to handle long conversations, and leverage more than 20+ APIs successfully, other models struggle with longer context or when increasing the number of APIs. We also report that the performance on chained function-calls is severely limited across the models. Overall, the top performing models on CONFETTI are Nova Pro (40.01%), Claude Sonnet v3.5 (35.46%) and Llama 3.1 405B (33.19%) followed by command-r-plus (31.18%) and Mistral-Large-2407 (30.07%).
title CONFETTI: Conversational Function-Calling Evaluation Through Turn-Level Interactions
topic Computation and Language
url https://arxiv.org/abs/2506.01859