FunctionChat-Bench: Comprehensive Evaluation of Language Models' Generative Capabilities in Korean Tool-use Dialogs

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Lee, Shinbok, Seo, Gaeun, Lee, Daniel, Ko, Byeongil, Jung, Sunghee, Shin, Myeongcheol
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915029449703424
author Lee, Shinbok
Seo, Gaeun
Lee, Daniel
Ko, Byeongil
Jung, Sunghee
Shin, Myeongcheol
author_facet Lee, Shinbok
Seo, Gaeun
Lee, Daniel
Ko, Byeongil
Jung, Sunghee
Shin, Myeongcheol
contents This study investigates language models' generative capabilities in tool-use dialogs. We categorize the models' outputs in tool-use dialogs into four distinct types: Tool Call, Answer Completion, Slot Question, and Relevance Detection, which serve as aspects for evaluation. We introduce FunctionChat-Bench, comprising 700 evaluation items and automated assessment programs. Using this benchmark, we evaluate several language models that support function calling. Our findings indicate that while language models may exhibit high accuracy in single-turn Tool Call scenarios, this does not necessarily translate to superior generative performance in multi-turn environments. We argue that the capabilities required for function calling extend beyond generating tool call messages; they must also effectively generate conversational messages that engage the user.
format Preprint
id arxiv_https___arxiv_org_abs_2411_14054
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle FunctionChat-Bench: Comprehensive Evaluation of Language Models' Generative Capabilities in Korean Tool-use Dialogs
Lee, Shinbok
Seo, Gaeun
Lee, Daniel
Ko, Byeongil
Jung, Sunghee
Shin, Myeongcheol
Computation and Language
Artificial Intelligence
This study investigates language models' generative capabilities in tool-use dialogs. We categorize the models' outputs in tool-use dialogs into four distinct types: Tool Call, Answer Completion, Slot Question, and Relevance Detection, which serve as aspects for evaluation. We introduce FunctionChat-Bench, comprising 700 evaluation items and automated assessment programs. Using this benchmark, we evaluate several language models that support function calling. Our findings indicate that while language models may exhibit high accuracy in single-turn Tool Call scenarios, this does not necessarily translate to superior generative performance in multi-turn environments. We argue that the capabilities required for function calling extend beyond generating tool call messages; they must also effectively generate conversational messages that engage the user.
title FunctionChat-Bench: Comprehensive Evaluation of Language Models' Generative Capabilities in Korean Tool-use Dialogs
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2411.14054