A Collection of Question Answering Datasets for Norwegian

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Mikhailov, Vladislav, Mæhlum, Petter, Langø, Victoria Ovedie Chruickshank, Velldal, Erik, Øvrelid, Lilja
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910790071615488
author Mikhailov, Vladislav
Mæhlum, Petter
Langø, Victoria Ovedie Chruickshank
Velldal, Erik
Øvrelid, Lilja
author_facet Mikhailov, Vladislav
Mæhlum, Petter
Langø, Victoria Ovedie Chruickshank
Velldal, Erik
Øvrelid, Lilja
contents This paper introduces a new suite of question answering datasets for Norwegian; NorOpenBookQA, NorCommonSenseQA, NorTruthfulQA, and NRK-Quiz-QA. The data covers a wide range of skills and knowledge domains, including world knowledge, commonsense reasoning, truthfulness, and knowledge about Norway. Covering both of the written standards of Norwegian - Bokmål and Nynorsk - our datasets comprise over 10k question-answer pairs, created by native speakers. We detail our dataset creation approach and present the results of evaluating 11 language models (LMs) in zero- and few-shot regimes. Most LMs perform better in Bokmål than Nynorsk, struggle most with commonsense reasoning, and are often untruthful in generating answers to questions. All our datasets and annotation materials are publicly available.
format Preprint
id arxiv_https___arxiv_org_abs_2501_11128
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Collection of Question Answering Datasets for Norwegian
Mikhailov, Vladislav
Mæhlum, Petter
Langø, Victoria Ovedie Chruickshank
Velldal, Erik
Øvrelid, Lilja
Computation and Language
Artificial Intelligence
This paper introduces a new suite of question answering datasets for Norwegian; NorOpenBookQA, NorCommonSenseQA, NorTruthfulQA, and NRK-Quiz-QA. The data covers a wide range of skills and knowledge domains, including world knowledge, commonsense reasoning, truthfulness, and knowledge about Norway. Covering both of the written standards of Norwegian - Bokmål and Nynorsk - our datasets comprise over 10k question-answer pairs, created by native speakers. We detail our dataset creation approach and present the results of evaluating 11 language models (LMs) in zero- and few-shot regimes. Most LMs perform better in Bokmål than Nynorsk, struggle most with commonsense reasoning, and are often untruthful in generating answers to questions. All our datasets and annotation materials are publicly available.
title A Collection of Question Answering Datasets for Norwegian
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2501.11128