Tagengo: A Multilingual Chat Dataset

Fuente: arXiv
Saved in:
Bibliographic Details
Main Author: Devine, Peter
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916254225268736
author Devine, Peter
author_facet Devine, Peter
contents Open source large language models (LLMs) have shown great improvements in recent times. However, many of these models are focused solely on popular spoken languages. We present a high quality dataset of more than 70k prompt-response pairs in 74 languages which consist of human generated prompts and synthetic responses. We use this dataset to train a state-of-the-art open source English LLM to chat multilingually. We evaluate our model on MT-Bench chat benchmarks in 6 languages, finding that our multilingual model outperforms previous state-of-the-art open source LLMs across each language. We further find that training on more multilingual data is beneficial to the performance in a chosen target language (Japanese) compared to simply training on only data in that language. These results indicate the necessity of training on large amounts of high quality multilingual data to make a more accessible LLM.
format Preprint
id arxiv_https___arxiv_org_abs_2405_12612
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Tagengo: A Multilingual Chat Dataset
Devine, Peter
Computation and Language
Artificial Intelligence
Machine Learning
Open source large language models (LLMs) have shown great improvements in recent times. However, many of these models are focused solely on popular spoken languages. We present a high quality dataset of more than 70k prompt-response pairs in 74 languages which consist of human generated prompts and synthetic responses. We use this dataset to train a state-of-the-art open source English LLM to chat multilingually. We evaluate our model on MT-Bench chat benchmarks in 6 languages, finding that our multilingual model outperforms previous state-of-the-art open source LLMs across each language. We further find that training on more multilingual data is beneficial to the performance in a chosen target language (Japanese) compared to simply training on only data in that language. These results indicate the necessity of training on large amounts of high quality multilingual data to make a more accessible LLM.
title Tagengo: A Multilingual Chat Dataset
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2405.12612