"Vorbeşti Româneşte?" A Recipe to Train Powerful Romanian LLMs with English Instructions

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Masala, Mihai, Ilie-Ablachim, Denis C., Dima, Alexandru, Corlatescu, Dragos, Zavelca, Miruna, Olaru, Ovio, Terian, Simina, Terian, Andrei, Leordeanu, Marius, Velicu, Horia, Popescu, Marius, Dascalu, Mihai, Rebedea, Traian
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916450154840064
author Masala, Mihai
Ilie-Ablachim, Denis C.
Dima, Alexandru
Corlatescu, Dragos
Zavelca, Miruna
Olaru, Ovio
Terian, Simina
Terian, Andrei
Leordeanu, Marius
Velicu, Horia
Popescu, Marius
Dascalu, Mihai
Rebedea, Traian
author_facet Masala, Mihai
Ilie-Ablachim, Denis C.
Dima, Alexandru
Corlatescu, Dragos
Zavelca, Miruna
Olaru, Ovio
Terian, Simina
Terian, Andrei
Leordeanu, Marius
Velicu, Horia
Popescu, Marius
Dascalu, Mihai
Rebedea, Traian
contents In recent years, Large Language Models (LLMs) have achieved almost human-like performance on various tasks. While some LLMs have been trained on multilingual data, most of the training data is in English; hence, their performance in English greatly exceeds other languages. To our knowledge, we are the first to collect and translate a large collection of texts, instructions, and benchmarks and train, evaluate, and release open-source LLMs tailored for Romanian. We evaluate our methods on four different categories, including academic benchmarks, MT-Bench (manually translated), and a professionally built historical, cultural, and social benchmark adapted to Romanian. We argue for the usefulness and high performance of RoLLMs by obtaining state-of-the-art results across the board. We publicly release all resources (i.e., data, training and evaluation code, models) to support and encourage research on Romanian LLMs while concurrently creating a generalizable recipe, adequate for other low or less-resourced languages.
format Preprint
id arxiv_https___arxiv_org_abs_2406_18266
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle "Vorbeşti Româneşte?" A Recipe to Train Powerful Romanian LLMs with English Instructions
Masala, Mihai
Ilie-Ablachim, Denis C.
Dima, Alexandru
Corlatescu, Dragos
Zavelca, Miruna
Olaru, Ovio
Terian, Simina
Terian, Andrei
Leordeanu, Marius
Velicu, Horia
Popescu, Marius
Dascalu, Mihai
Rebedea, Traian
Computation and Language
In recent years, Large Language Models (LLMs) have achieved almost human-like performance on various tasks. While some LLMs have been trained on multilingual data, most of the training data is in English; hence, their performance in English greatly exceeds other languages. To our knowledge, we are the first to collect and translate a large collection of texts, instructions, and benchmarks and train, evaluate, and release open-source LLMs tailored for Romanian. We evaluate our methods on four different categories, including academic benchmarks, MT-Bench (manually translated), and a professionally built historical, cultural, and social benchmark adapted to Romanian. We argue for the usefulness and high performance of RoLLMs by obtaining state-of-the-art results across the board. We publicly release all resources (i.e., data, training and evaluation code, models) to support and encourage research on Romanian LLMs while concurrently creating a generalizable recipe, adequate for other low or less-resourced languages.
title "Vorbeşti Româneşte?" A Recipe to Train Powerful Romanian LLMs with English Instructions
topic Computation and Language
url https://arxiv.org/abs/2406.18266