Question: How do Large Language Models perform on the Question Answering tasks? Answer:

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fischer, Kevin, Fürst, Darren, Steindl, Sebastian, Lindner, Jakob, Schäfer, Ulrich
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913615821406208
author Fischer, Kevin
Fürst, Darren
Steindl, Sebastian
Lindner, Jakob
Schäfer, Ulrich
author_facet Fischer, Kevin
Fürst, Darren
Steindl, Sebastian
Lindner, Jakob
Schäfer, Ulrich
contents Large Language Models (LLMs) have been showing promising results for various NLP-tasks without the explicit need to be trained for these tasks by using few-shot or zero-shot prompting techniques. A common NLP-task is question-answering (QA). In this study, we propose a comprehensive performance comparison between smaller fine-tuned models and out-of-the-box instruction-following LLMs on the Stanford Question Answering Dataset 2.0 (SQuAD2), specifically when using a single-inference prompting technique. Since the dataset contains unanswerable questions, previous work used a double inference method. We propose a prompting style which aims to elicit the same ability without the need for double inference, saving compute time and resources. Furthermore, we investigate their generalization capabilities by comparing their performance on similar but different QA datasets, without fine-tuning neither model, emulating real-world uses where the context and questions asked may differ from the original training distribution, for example swapping Wikipedia for news articles. Our results show that smaller, fine-tuned models outperform current State-Of-The-Art (SOTA) LLMs on the fine-tuned task, but recent SOTA models are able to close this gap on the out-of-distribution test and even outperform the fine-tuned models on 3 of the 5 tested QA datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2412_12893
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Question: How do Large Language Models perform on the Question Answering tasks? Answer:
Fischer, Kevin
Fürst, Darren
Steindl, Sebastian
Lindner, Jakob
Schäfer, Ulrich
Computation and Language
Large Language Models (LLMs) have been showing promising results for various NLP-tasks without the explicit need to be trained for these tasks by using few-shot or zero-shot prompting techniques. A common NLP-task is question-answering (QA). In this study, we propose a comprehensive performance comparison between smaller fine-tuned models and out-of-the-box instruction-following LLMs on the Stanford Question Answering Dataset 2.0 (SQuAD2), specifically when using a single-inference prompting technique. Since the dataset contains unanswerable questions, previous work used a double inference method. We propose a prompting style which aims to elicit the same ability without the need for double inference, saving compute time and resources. Furthermore, we investigate their generalization capabilities by comparing their performance on similar but different QA datasets, without fine-tuning neither model, emulating real-world uses where the context and questions asked may differ from the original training distribution, for example swapping Wikipedia for news articles. Our results show that smaller, fine-tuned models outperform current State-Of-The-Art (SOTA) LLMs on the fine-tuned task, but recent SOTA models are able to close this gap on the out-of-distribution test and even outperform the fine-tuned models on 3 of the 5 tested QA datasets.
title Question: How do Large Language Models perform on the Question Answering tasks? Answer:
topic Computation and Language
url https://arxiv.org/abs/2412.12893