Fineweb-Edu-Ar: Machine-translated Corpus to Support Arabic Small Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Alrashed, Sultan, Khizbullin, Dmitrii, Pugh, David R.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917833169960960
author Alrashed, Sultan
Khizbullin, Dmitrii
Pugh, David R.
author_facet Alrashed, Sultan
Khizbullin, Dmitrii
Pugh, David R.
contents As large language models (LLMs) grow and develop, so do their data demands. This is especially true for multilingual LLMs, where the scarcity of high-quality and readily available data online has led to a multitude of synthetic dataset generation approaches. A key technique in this space is machine translation (MT), where high-quality English text is adapted to a target, comparatively low-resource language. This report introduces FineWeb-Edu-Ar, a machine-translated version of the exceedingly popular (deduplicated) FineWeb-Edu dataset from HuggingFace. To the best of our knowledge, FineWeb-Edu-Ar is the largest publicly available machine-translated Arabic dataset out there, with its size of 202B tokens of an Arabic-trained tokenizer.
format Preprint
id arxiv_https___arxiv_org_abs_2411_06402
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Fineweb-Edu-Ar: Machine-translated Corpus to Support Arabic Small Language Models
Alrashed, Sultan
Khizbullin, Dmitrii
Pugh, David R.
Computation and Language
Artificial Intelligence
As large language models (LLMs) grow and develop, so do their data demands. This is especially true for multilingual LLMs, where the scarcity of high-quality and readily available data online has led to a multitude of synthetic dataset generation approaches. A key technique in this space is machine translation (MT), where high-quality English text is adapted to a target, comparatively low-resource language. This report introduces FineWeb-Edu-Ar, a machine-translated version of the exceedingly popular (deduplicated) FineWeb-Edu dataset from HuggingFace. To the best of our knowledge, FineWeb-Edu-Ar is the largest publicly available machine-translated Arabic dataset out there, with its size of 202B tokens of an Arabic-trained tokenizer.
title Fineweb-Edu-Ar: Machine-translated Corpus to Support Arabic Small Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2411.06402