New Textual Corpora for Serbian Language Modeling

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Škorić, Mihailo, Janković, Nikola
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911877181734912
author Škorić, Mihailo
Janković, Nikola
author_facet Škorić, Mihailo
Janković, Nikola
contents This paper will present textual corpora for Serbian (and Serbo-Croatian), usable for the training of large language models and publicly available at one of the several notable online repositories. Each corpus will be classified using multiple methods and its characteristics will be detailed. Additionally, the paper will introduce three new corpora: a new umbrella web corpus of Serbo-Croatian, a new high-quality corpus based on the doctoral dissertations stored within National Repository of Doctoral Dissertations from all Universities in Serbia, and a parallel corpus of abstract translation from the same source. The uniqueness of both old and new corpora will be accessed via frequency-based stylometric methods, and the results will be briefly discussed.
format Preprint
id arxiv_https___arxiv_org_abs_2405_09250
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle New Textual Corpora for Serbian Language Modeling
Škorić, Mihailo
Janković, Nikola
Computation and Language
This paper will present textual corpora for Serbian (and Serbo-Croatian), usable for the training of large language models and publicly available at one of the several notable online repositories. Each corpus will be classified using multiple methods and its characteristics will be detailed. Additionally, the paper will introduce three new corpora: a new umbrella web corpus of Serbo-Croatian, a new high-quality corpus based on the doctoral dissertations stored within National Repository of Doctoral Dissertations from all Universities in Serbia, and a parallel corpus of abstract translation from the same source. The uniqueness of both old and new corpora will be accessed via frequency-based stylometric methods, and the results will be briefly discussed.
title New Textual Corpora for Serbian Language Modeling
topic Computation and Language
url https://arxiv.org/abs/2405.09250