SWEb: A Large Web Dataset for the Scandinavian Languages

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Norlund, Tobias, Isbister, Tim, Gyllensten, Amaru Cuba, Santos, Paul Dos, Petrelli, Danila, Ekgren, Ariel, Sahlgren, Magnus
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914965847277568
author Norlund, Tobias
Isbister, Tim
Gyllensten, Amaru Cuba
Santos, Paul Dos
Petrelli, Danila
Ekgren, Ariel
Sahlgren, Magnus
author_facet Norlund, Tobias
Isbister, Tim
Gyllensten, Amaru Cuba
Santos, Paul Dos
Petrelli, Danila
Ekgren, Ariel
Sahlgren, Magnus
contents This paper presents the hitherto largest pretraining dataset for the Scandinavian languages: the Scandinavian WEb (SWEb), comprising over one trillion tokens. The paper details the collection and processing pipeline, and introduces a novel model-based text extractor that significantly reduces complexity in comparison with rule-based approaches. We also introduce a new cloze-style benchmark for evaluating language models in Swedish, and use this test to compare models trained on the SWEb data to models trained on FineWeb, with competitive results. All data, models and code are shared openly.
format Preprint
id arxiv_https___arxiv_org_abs_2410_04456
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SWEb: A Large Web Dataset for the Scandinavian Languages
Norlund, Tobias
Isbister, Tim
Gyllensten, Amaru Cuba
Santos, Paul Dos
Petrelli, Danila
Ekgren, Ariel
Sahlgren, Magnus
Computation and Language
This paper presents the hitherto largest pretraining dataset for the Scandinavian languages: the Scandinavian WEb (SWEb), comprising over one trillion tokens. The paper details the collection and processing pipeline, and introduces a novel model-based text extractor that significantly reduces complexity in comparison with rule-based approaches. We also introduce a new cloze-style benchmark for evaluating language models in Swedish, and use this test to compare models trained on the SWEb data to models trained on FineWeb, with competitive results. All data, models and code are shared openly.
title SWEb: A Large Web Dataset for the Scandinavian Languages
topic Computation and Language
url https://arxiv.org/abs/2410.04456