CATfOOD: Counterfactual Augmented Training for Improving Out-of-Domain Performance and Calibration

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Sachdeva, Rachneet, Tutek, Martin, Gurevych, Iryna
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866929241518505984
author Sachdeva, Rachneet
Tutek, Martin
Gurevych, Iryna
author_facet Sachdeva, Rachneet
Tutek, Martin
Gurevych, Iryna
contents In recent years, large language models (LLMs) have shown remarkable capabilities at scale, particularly at generating text conditioned on a prompt. In our work, we investigate the use of LLMs to augment training data of small language models~(SLMs) with automatically generated counterfactual~(CF) instances -- i.e. minimally altered inputs -- in order to improve out-of-domain~(OOD) performance of SLMs in the extractive question answering~(QA) setup. We show that, across various LLM generators, such data augmentation consistently enhances OOD performance and improves model calibration for both confidence-based and rationale-augmented calibrator models. Furthermore, these performance improvements correlate with higher diversity of CF instances in terms of their surface form and semantic content. Finally, we show that CF augmented models which are easier to calibrate also exhibit much lower entropy when assigning importance, indicating that rationale-augmented calibrators prefer concise explanations.
format Preprint
id arxiv_https___arxiv_org_abs_2309_07822
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle CATfOOD: Counterfactual Augmented Training for Improving Out-of-Domain Performance and Calibration
Sachdeva, Rachneet
Tutek, Martin
Gurevych, Iryna
Computation and Language
In recent years, large language models (LLMs) have shown remarkable capabilities at scale, particularly at generating text conditioned on a prompt. In our work, we investigate the use of LLMs to augment training data of small language models~(SLMs) with automatically generated counterfactual~(CF) instances -- i.e. minimally altered inputs -- in order to improve out-of-domain~(OOD) performance of SLMs in the extractive question answering~(QA) setup. We show that, across various LLM generators, such data augmentation consistently enhances OOD performance and improves model calibration for both confidence-based and rationale-augmented calibrator models. Furthermore, these performance improvements correlate with higher diversity of CF instances in terms of their surface form and semantic content. Finally, we show that CF augmented models which are easier to calibrate also exhibit much lower entropy when assigning importance, indicating that rationale-augmented calibrators prefer concise explanations.
title CATfOOD: Counterfactual Augmented Training for Improving Out-of-Domain Performance and Calibration
topic Computation and Language
url https://arxiv.org/abs/2309.07822