"Vorbeşti Româneşte?" A Recipe to Train Powerful Romanian LLMs with English Instructions
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866916450154840064 |
|---|---|
| author | Masala, Mihai Ilie-Ablachim, Denis C. Dima, Alexandru Corlatescu, Dragos Zavelca, Miruna Olaru, Ovio Terian, Simina Terian, Andrei Leordeanu, Marius Velicu, Horia Popescu, Marius Dascalu, Mihai Rebedea, Traian |
| author_facet | Masala, Mihai Ilie-Ablachim, Denis C. Dima, Alexandru Corlatescu, Dragos Zavelca, Miruna Olaru, Ovio Terian, Simina Terian, Andrei Leordeanu, Marius Velicu, Horia Popescu, Marius Dascalu, Mihai Rebedea, Traian |
| contents | In recent years, Large Language Models (LLMs) have achieved almost human-like performance on various tasks. While some LLMs have been trained on multilingual data, most of the training data is in English; hence, their performance in English greatly exceeds other languages. To our knowledge, we are the first to collect and translate a large collection of texts, instructions, and benchmarks and train, evaluate, and release open-source LLMs tailored for Romanian. We evaluate our methods on four different categories, including academic benchmarks, MT-Bench (manually translated), and a professionally built historical, cultural, and social benchmark adapted to Romanian. We argue for the usefulness and high performance of RoLLMs by obtaining state-of-the-art results across the board. We publicly release all resources (i.e., data, training and evaluation code, models) to support and encourage research on Romanian LLMs while concurrently creating a generalizable recipe, adequate for other low or less-resourced languages. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2406_18266 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | "Vorbeşti Româneşte?" A Recipe to Train Powerful Romanian LLMs with English Instructions Masala, Mihai Ilie-Ablachim, Denis C. Dima, Alexandru Corlatescu, Dragos Zavelca, Miruna Olaru, Ovio Terian, Simina Terian, Andrei Leordeanu, Marius Velicu, Horia Popescu, Marius Dascalu, Mihai Rebedea, Traian Computation and Language In recent years, Large Language Models (LLMs) have achieved almost human-like performance on various tasks. While some LLMs have been trained on multilingual data, most of the training data is in English; hence, their performance in English greatly exceeds other languages. To our knowledge, we are the first to collect and translate a large collection of texts, instructions, and benchmarks and train, evaluate, and release open-source LLMs tailored for Romanian. We evaluate our methods on four different categories, including academic benchmarks, MT-Bench (manually translated), and a professionally built historical, cultural, and social benchmark adapted to Romanian. We argue for the usefulness and high performance of RoLLMs by obtaining state-of-the-art results across the board. We publicly release all resources (i.e., data, training and evaluation code, models) to support and encourage research on Romanian LLMs while concurrently creating a generalizable recipe, adequate for other low or less-resourced languages. |
| title | "Vorbeşti Româneşte?" A Recipe to Train Powerful Romanian LLMs with English Instructions |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2406.18266 |