Matina: A Large-Scale 73B Token Persian Text Corpus

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Hosseinbeigi, Sara Bourbour, Taherinezhad, Fatemeh, Faili, Heshaam, Baghbani, Hamed, Nadi, Fatemeh, Amiri, Mostafa
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909491143901184
author Hosseinbeigi, Sara Bourbour
Taherinezhad, Fatemeh
Faili, Heshaam
Baghbani, Hamed
Nadi, Fatemeh
Amiri, Mostafa
author_facet Hosseinbeigi, Sara Bourbour
Taherinezhad, Fatemeh
Faili, Heshaam
Baghbani, Hamed
Nadi, Fatemeh
Amiri, Mostafa
contents Text corpora are essential for training models used in tasks like summarization, translation, and large language models (LLMs). While various efforts have been made to collect monolingual and multilingual datasets in many languages, Persian has often been underrepresented due to limited resources for data collection and preprocessing. Existing Persian datasets are typically small and lack content diversity, consisting mainly of weblogs and news articles. This shortage of high-quality, varied data has slowed the development of NLP models and open-source LLMs for Persian. Since model performance depends heavily on the quality of training data, we address this gap by introducing the Matina corpus, a new Persian dataset of 72.9B tokens, carefully preprocessed and deduplicated to ensure high data quality. We further assess its effectiveness by training and evaluating transformer-based models on key NLP tasks. Both the dataset and preprocessing codes are publicly available, enabling researchers to build on and improve this resource for future Persian NLP advancements.
format Preprint
id arxiv_https___arxiv_org_abs_2502_09188
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Matina: A Large-Scale 73B Token Persian Text Corpus
Hosseinbeigi, Sara Bourbour
Taherinezhad, Fatemeh
Faili, Heshaam
Baghbani, Hamed
Nadi, Fatemeh
Amiri, Mostafa
Computation and Language
Artificial Intelligence
Text corpora are essential for training models used in tasks like summarization, translation, and large language models (LLMs). While various efforts have been made to collect monolingual and multilingual datasets in many languages, Persian has often been underrepresented due to limited resources for data collection and preprocessing. Existing Persian datasets are typically small and lack content diversity, consisting mainly of weblogs and news articles. This shortage of high-quality, varied data has slowed the development of NLP models and open-source LLMs for Persian. Since model performance depends heavily on the quality of training data, we address this gap by introducing the Matina corpus, a new Persian dataset of 72.9B tokens, carefully preprocessed and deduplicated to ensure high data quality. We further assess its effectiveness by training and evaluating transformer-based models on key NLP tasks. Both the dataset and preprocessing codes are publicly available, enabling researchers to build on and improve this resource for future Persian NLP advancements.
title Matina: A Large-Scale 73B Token Persian Text Corpus
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2502.09188