Generative Data Augmentation using LLMs improves Distributional Robustness in Question Answering

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chowdhury, Arijit Ghosh, Chadha, Aman
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914671763652608
author Chowdhury, Arijit Ghosh
Chadha, Aman
author_facet Chowdhury, Arijit Ghosh
Chadha, Aman
contents Robustness in Natural Language Processing continues to be a pertinent issue, where state of the art models under-perform under naturally shifted distributions. In the context of Question Answering, work on domain adaptation methods continues to be a growing body of research. However, very little attention has been given to the notion of domain generalization under natural distribution shifts, where the target domain is unknown. With drastic improvements in the quality and access to generative models, we answer the question: How do generated datasets influence the performance of QA models under natural distribution shifts? We perform experiments on 4 different datasets under varying amounts of distribution shift, and analyze how "in-the-wild" generation can help achieve domain generalization. We take a two-step generation approach, generating both contexts and QA pairs to augment existing datasets. Through our experiments, we demonstrate how augmenting reading comprehension datasets with generated data leads to better robustness towards natural distribution shifts.
format Preprint
id arxiv_https___arxiv_org_abs_2309_06358
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Generative Data Augmentation using LLMs improves Distributional Robustness in Question Answering
Chowdhury, Arijit Ghosh
Chadha, Aman
Computation and Language
Artificial Intelligence
Robustness in Natural Language Processing continues to be a pertinent issue, where state of the art models under-perform under naturally shifted distributions. In the context of Question Answering, work on domain adaptation methods continues to be a growing body of research. However, very little attention has been given to the notion of domain generalization under natural distribution shifts, where the target domain is unknown. With drastic improvements in the quality and access to generative models, we answer the question: How do generated datasets influence the performance of QA models under natural distribution shifts? We perform experiments on 4 different datasets under varying amounts of distribution shift, and analyze how "in-the-wild" generation can help achieve domain generalization. We take a two-step generation approach, generating both contexts and QA pairs to augment existing datasets. Through our experiments, we demonstrate how augmenting reading comprehension datasets with generated data leads to better robustness towards natural distribution shifts.
title Generative Data Augmentation using LLMs improves Distributional Robustness in Question Answering
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2309.06358