NewsQs: Multi-Source Question Generation for the Inquiring Mind
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911918908768256 |
|---|---|
| author | Hwang, Alyssa Dixit, Kalpit Ballesteros, Miguel Benajiba, Yassine Castelli, Vittorio Dreyer, Markus Bansal, Mohit McKeown, Kathleen |
| author_facet | Hwang, Alyssa Dixit, Kalpit Ballesteros, Miguel Benajiba, Yassine Castelli, Vittorio Dreyer, Markus Bansal, Mohit McKeown, Kathleen |
| contents | We present NewsQs (news-cues), a dataset that provides question-answer pairs for multiple news documents. To create NewsQs, we augment a traditional multi-document summarization dataset with questions automatically generated by a T5-Large model fine-tuned on FAQ-style news articles from the News On the Web corpus. We show that fine-tuning a model with control codes produces questions that are judged acceptable more often than the same model without them as measured through human evaluation. We use a QNLI model with high correlation with human annotations to filter our data. We release our final dataset of high-quality questions, answers, and document clusters as a resource for future work in query-based multi-document summarization. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2402_18479 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | NewsQs: Multi-Source Question Generation for the Inquiring Mind Hwang, Alyssa Dixit, Kalpit Ballesteros, Miguel Benajiba, Yassine Castelli, Vittorio Dreyer, Markus Bansal, Mohit McKeown, Kathleen Computation and Language We present NewsQs (news-cues), a dataset that provides question-answer pairs for multiple news documents. To create NewsQs, we augment a traditional multi-document summarization dataset with questions automatically generated by a T5-Large model fine-tuned on FAQ-style news articles from the News On the Web corpus. We show that fine-tuning a model with control codes produces questions that are judged acceptable more often than the same model without them as measured through human evaluation. We use a QNLI model with high correlation with human annotations to filter our data. We release our final dataset of high-quality questions, answers, and document clusters as a resource for future work in query-based multi-document summarization. |
| title | NewsQs: Multi-Source Question Generation for the Inquiring Mind |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2402.18479 |