NewsQs: Multi-Source Question Generation for the Inquiring Mind

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hwang, Alyssa, Dixit, Kalpit, Ballesteros, Miguel, Benajiba, Yassine, Castelli, Vittorio, Dreyer, Markus, Bansal, Mohit, McKeown, Kathleen
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911918908768256
author Hwang, Alyssa
Dixit, Kalpit
Ballesteros, Miguel
Benajiba, Yassine
Castelli, Vittorio
Dreyer, Markus
Bansal, Mohit
McKeown, Kathleen
author_facet Hwang, Alyssa
Dixit, Kalpit
Ballesteros, Miguel
Benajiba, Yassine
Castelli, Vittorio
Dreyer, Markus
Bansal, Mohit
McKeown, Kathleen
contents We present NewsQs (news-cues), a dataset that provides question-answer pairs for multiple news documents. To create NewsQs, we augment a traditional multi-document summarization dataset with questions automatically generated by a T5-Large model fine-tuned on FAQ-style news articles from the News On the Web corpus. We show that fine-tuning a model with control codes produces questions that are judged acceptable more often than the same model without them as measured through human evaluation. We use a QNLI model with high correlation with human annotations to filter our data. We release our final dataset of high-quality questions, answers, and document clusters as a resource for future work in query-based multi-document summarization.
format Preprint
id arxiv_https___arxiv_org_abs_2402_18479
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle NewsQs: Multi-Source Question Generation for the Inquiring Mind
Hwang, Alyssa
Dixit, Kalpit
Ballesteros, Miguel
Benajiba, Yassine
Castelli, Vittorio
Dreyer, Markus
Bansal, Mohit
McKeown, Kathleen
Computation and Language
We present NewsQs (news-cues), a dataset that provides question-answer pairs for multiple news documents. To create NewsQs, we augment a traditional multi-document summarization dataset with questions automatically generated by a T5-Large model fine-tuned on FAQ-style news articles from the News On the Web corpus. We show that fine-tuning a model with control codes produces questions that are judged acceptable more often than the same model without them as measured through human evaluation. We use a QNLI model with high correlation with human annotations to filter our data. We release our final dataset of high-quality questions, answers, and document clusters as a resource for future work in query-based multi-document summarization.
title NewsQs: Multi-Source Question Generation for the Inquiring Mind
topic Computation and Language
url https://arxiv.org/abs/2402.18479