KIWI: A Dataset of Knowledge-Intensive Writing Instructions for Answering Research Questions

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Xu, Fangyuan, Lo, Kyle, Soldaini, Luca, Kuehl, Bailey, Choi, Eunsol, Wadden, David
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866929266359271424
author Xu, Fangyuan
Lo, Kyle
Soldaini, Luca
Kuehl, Bailey
Choi, Eunsol
Wadden, David
author_facet Xu, Fangyuan
Lo, Kyle
Soldaini, Luca
Kuehl, Bailey
Choi, Eunsol
Wadden, David
contents Large language models (LLMs) adapted to follow user instructions are now widely deployed as conversational agents. In this work, we examine one increasingly common instruction-following task: providing writing assistance to compose a long-form answer. To evaluate the capabilities of current LLMs on this task, we construct KIWI, a dataset of knowledge-intensive writing instructions in the scientific domain. Given a research question, an initial model-generated answer and a set of relevant papers, an expert annotator iteratively issues instructions for the model to revise and improve its answer. We collect 1,260 interaction turns from 234 interaction sessions with three state-of-the-art LLMs. Each turn includes a user instruction, a model response, and a human evaluation of the model response. Through a detailed analysis of the collected responses, we find that all models struggle to incorporate new information into an existing answer, and to perform precise and unambiguous edits. Further, we find that models struggle to judge whether their outputs successfully followed user instructions, with accuracy at least 10 points short of human agreement. Our findings indicate that KIWI will be a valuable resource to measure progress and improve LLMs' instruction-following capabilities for knowledge intensive writing tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2403_03866
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle KIWI: A Dataset of Knowledge-Intensive Writing Instructions for Answering Research Questions
Xu, Fangyuan
Lo, Kyle
Soldaini, Luca
Kuehl, Bailey
Choi, Eunsol
Wadden, David
Computation and Language
Large language models (LLMs) adapted to follow user instructions are now widely deployed as conversational agents. In this work, we examine one increasingly common instruction-following task: providing writing assistance to compose a long-form answer. To evaluate the capabilities of current LLMs on this task, we construct KIWI, a dataset of knowledge-intensive writing instructions in the scientific domain. Given a research question, an initial model-generated answer and a set of relevant papers, an expert annotator iteratively issues instructions for the model to revise and improve its answer. We collect 1,260 interaction turns from 234 interaction sessions with three state-of-the-art LLMs. Each turn includes a user instruction, a model response, and a human evaluation of the model response. Through a detailed analysis of the collected responses, we find that all models struggle to incorporate new information into an existing answer, and to perform precise and unambiguous edits. Further, we find that models struggle to judge whether their outputs successfully followed user instructions, with accuracy at least 10 points short of human agreement. Our findings indicate that KIWI will be a valuable resource to measure progress and improve LLMs' instruction-following capabilities for knowledge intensive writing tasks.
title KIWI: A Dataset of Knowledge-Intensive Writing Instructions for Answering Research Questions
topic Computation and Language
url https://arxiv.org/abs/2403.03866