Evaluation of Instruction-Following Ability for Large Language Models on Story-Ending Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hida, Rem, Ohmura, Junki, Sekiya, Toshiyuki
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929397488943104
author Hida, Rem
Ohmura, Junki
Sekiya, Toshiyuki
author_facet Hida, Rem
Ohmura, Junki
Sekiya, Toshiyuki
contents Instruction-tuned Large Language Models (LLMs) have achieved remarkable performance across various benchmark tasks. While providing instructions to LLMs for guiding their generations is user-friendly, assessing their instruction-following capabilities is still unclarified due to a lack of evaluation metrics. In this paper, we focus on evaluating the instruction-following ability of LLMs in the context of story-ending generation, which requires diverse and context-specific instructions. We propose an automatic evaluation pipeline that utilizes a machine reading comprehension (MRC) model to determine whether the generated story-ending reflects instruction. Our findings demonstrate that our proposed metric aligns with human evaluation. Furthermore, our experiments confirm that recent open-source LLMs can achieve instruction-following performance close to GPT-3.5, as assessed through automatic evaluation.
format Preprint
id arxiv_https___arxiv_org_abs_2406_16356
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Evaluation of Instruction-Following Ability for Large Language Models on Story-Ending Generation
Hida, Rem
Ohmura, Junki
Sekiya, Toshiyuki
Computation and Language
Instruction-tuned Large Language Models (LLMs) have achieved remarkable performance across various benchmark tasks. While providing instructions to LLMs for guiding their generations is user-friendly, assessing their instruction-following capabilities is still unclarified due to a lack of evaluation metrics. In this paper, we focus on evaluating the instruction-following ability of LLMs in the context of story-ending generation, which requires diverse and context-specific instructions. We propose an automatic evaluation pipeline that utilizes a machine reading comprehension (MRC) model to determine whether the generated story-ending reflects instruction. Our findings demonstrate that our proposed metric aligns with human evaluation. Furthermore, our experiments confirm that recent open-source LLMs can achieve instruction-following performance close to GPT-3.5, as assessed through automatic evaluation.
title Evaluation of Instruction-Following Ability for Large Language Models on Story-Ending Generation
topic Computation and Language
url https://arxiv.org/abs/2406.16356