Find the Intention of Instruction: Comprehensive Evaluation of Instruction Understanding for Large Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Moon, Hyeonseok, Seo, Jaehyung, Lee, Seungyoon, Park, Chanjun, Lim, Heuiseok
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915117271089152
author Moon, Hyeonseok
Seo, Jaehyung
Lee, Seungyoon
Park, Chanjun
Lim, Heuiseok
author_facet Moon, Hyeonseok
Seo, Jaehyung
Lee, Seungyoon
Park, Chanjun
Lim, Heuiseok
contents One of the key strengths of Large Language Models (LLMs) is their ability to interact with humans by generating appropriate responses to given instructions. This ability, known as instruction-following capability, has established a foundation for the use of LLMs across various fields and serves as a crucial metric for evaluating their performance. While numerous evaluation benchmarks have been developed, most focus solely on clear and coherent instructions. However, we have noted that LLMs can become easily distracted by instruction-formatted statements, which may lead to an oversight of their instruction comprehension skills. To address this issue, we introduce the Intention of Instruction (IoInst) benchmark. This benchmark evaluates LLMs' capacity to remain focused and understand instructions without being misled by extraneous instructions. The primary objective of this benchmark is to identify the appropriate instruction that accurately guides the generation of a given context. Our findings suggest that even recently introduced state-of-the-art models still lack instruction understanding capability. Along with the proposition of IoInst in this study, we also present broad analyses of the several strategies potentially applicable to IoInst.
format Preprint
id arxiv_https___arxiv_org_abs_2412_19450
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Find the Intention of Instruction: Comprehensive Evaluation of Instruction Understanding for Large Language Models
Moon, Hyeonseok
Seo, Jaehyung
Lee, Seungyoon
Park, Chanjun
Lim, Heuiseok
Artificial Intelligence
One of the key strengths of Large Language Models (LLMs) is their ability to interact with humans by generating appropriate responses to given instructions. This ability, known as instruction-following capability, has established a foundation for the use of LLMs across various fields and serves as a crucial metric for evaluating their performance. While numerous evaluation benchmarks have been developed, most focus solely on clear and coherent instructions. However, we have noted that LLMs can become easily distracted by instruction-formatted statements, which may lead to an oversight of their instruction comprehension skills. To address this issue, we introduce the Intention of Instruction (IoInst) benchmark. This benchmark evaluates LLMs' capacity to remain focused and understand instructions without being misled by extraneous instructions. The primary objective of this benchmark is to identify the appropriate instruction that accurately guides the generation of a given context. Our findings suggest that even recently introduced state-of-the-art models still lack instruction understanding capability. Along with the proposition of IoInst in this study, we also present broad analyses of the several strategies potentially applicable to IoInst.
title Find the Intention of Instruction: Comprehensive Evaluation of Instruction Understanding for Large Language Models
topic Artificial Intelligence
url https://arxiv.org/abs/2412.19450