Small Languages, Big Models: A Study of Continual Training on Languages of Norway
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866916594661195776 |
|---|---|
| author | Samuel, David Mikhailov, Vladislav Velldal, Erik Øvrelid, Lilja Charpentier, Lucas Georges Gabriel Kutuzov, Andrey Oepen, Stephan |
| author_facet | Samuel, David Mikhailov, Vladislav Velldal, Erik Øvrelid, Lilja Charpentier, Lucas Georges Gabriel Kutuzov, Andrey Oepen, Stephan |
| contents | Training large language models requires vast amounts of data, posing a challenge for less widely spoken languages like Norwegian and even more so for truly low-resource languages like Northern Sámi. To address this issue, we present a novel three-stage continual training approach that substantially improves the downstream performance together with the inference efficiency for the target languages. Based on our findings, we train, evaluate, and openly release a new generative language model for Norwegian Bokmål, Nynorsk, and Northern Sámi with 11.4 billion parameters: NorMistral-11B. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2412_06484 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Small Languages, Big Models: A Study of Continual Training on Languages of Norway Samuel, David Mikhailov, Vladislav Velldal, Erik Øvrelid, Lilja Charpentier, Lucas Georges Gabriel Kutuzov, Andrey Oepen, Stephan Computation and Language Training large language models requires vast amounts of data, posing a challenge for less widely spoken languages like Norwegian and even more so for truly low-resource languages like Northern Sámi. To address this issue, we present a novel three-stage continual training approach that substantially improves the downstream performance together with the inference efficiency for the target languages. Based on our findings, we train, evaluate, and openly release a new generative language model for Norwegian Bokmål, Nynorsk, and Northern Sámi with 11.4 billion parameters: NorMistral-11B. |
| title | Small Languages, Big Models: A Study of Continual Training on Languages of Norway |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2412.06484 |