Small Languages, Big Models: A Study of Continual Training on Languages of Norway

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Samuel, David, Mikhailov, Vladislav, Velldal, Erik, Øvrelid, Lilja, Charpentier, Lucas Georges Gabriel, Kutuzov, Andrey, Oepen, Stephan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916594661195776
author Samuel, David
Mikhailov, Vladislav
Velldal, Erik
Øvrelid, Lilja
Charpentier, Lucas Georges Gabriel
Kutuzov, Andrey
Oepen, Stephan
author_facet Samuel, David
Mikhailov, Vladislav
Velldal, Erik
Øvrelid, Lilja
Charpentier, Lucas Georges Gabriel
Kutuzov, Andrey
Oepen, Stephan
contents Training large language models requires vast amounts of data, posing a challenge for less widely spoken languages like Norwegian and even more so for truly low-resource languages like Northern Sámi. To address this issue, we present a novel three-stage continual training approach that substantially improves the downstream performance together with the inference efficiency for the target languages. Based on our findings, we train, evaluate, and openly release a new generative language model for Norwegian Bokmål, Nynorsk, and Northern Sámi with 11.4 billion parameters: NorMistral-11B.
format Preprint
id arxiv_https___arxiv_org_abs_2412_06484
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Small Languages, Big Models: A Study of Continual Training on Languages of Norway
Samuel, David
Mikhailov, Vladislav
Velldal, Erik
Øvrelid, Lilja
Charpentier, Lucas Georges Gabriel
Kutuzov, Andrey
Oepen, Stephan
Computation and Language
Training large language models requires vast amounts of data, posing a challenge for less widely spoken languages like Norwegian and even more so for truly low-resource languages like Northern Sámi. To address this issue, we present a novel three-stage continual training approach that substantially improves the downstream performance together with the inference efficiency for the target languages. Based on our findings, we train, evaluate, and openly release a new generative language model for Norwegian Bokmål, Nynorsk, and Northern Sámi with 11.4 billion parameters: NorMistral-11B.
title Small Languages, Big Models: A Study of Continual Training on Languages of Norway
topic Computation and Language
url https://arxiv.org/abs/2412.06484