Open foundation models for Azerbaijani language

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Isbarov, Jafar, Huseynova, Kavsar, Mammadov, Elvin, Hajili, Mammad, Ataman, Duygu
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929465133629440
author Isbarov, Jafar
Huseynova, Kavsar
Mammadov, Elvin
Hajili, Mammad
Ataman, Duygu
author_facet Isbarov, Jafar
Huseynova, Kavsar
Mammadov, Elvin
Hajili, Mammad
Ataman, Duygu
contents The emergence of multilingual large language models has enabled the development of language understanding and generation systems in Azerbaijani. However, most of the production-grade systems rely on cloud solutions, such as GPT-4. While there have been several attempts to develop open foundation models for Azerbaijani, these works have not found their way into common use due to a lack of systemic benchmarking. This paper encompasses several lines of work that promote open-source foundation models for Azerbaijani. We introduce (1) a large text corpus for Azerbaijani, (2) a family of encoder-only language models trained on this dataset, (3) labeled datasets for evaluating these models, and (4) extensive evaluation that covers all major open-source models with Azerbaijani support.
format Preprint
id arxiv_https___arxiv_org_abs_2407_02337
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Open foundation models for Azerbaijani language
Isbarov, Jafar
Huseynova, Kavsar
Mammadov, Elvin
Hajili, Mammad
Ataman, Duygu
Computation and Language
The emergence of multilingual large language models has enabled the development of language understanding and generation systems in Azerbaijani. However, most of the production-grade systems rely on cloud solutions, such as GPT-4. While there have been several attempts to develop open foundation models for Azerbaijani, these works have not found their way into common use due to a lack of systemic benchmarking. This paper encompasses several lines of work that promote open-source foundation models for Azerbaijani. We introduce (1) a large text corpus for Azerbaijani, (2) a family of encoder-only language models trained on this dataset, (3) labeled datasets for evaluating these models, and (4) extensive evaluation that covers all major open-source models with Azerbaijani support.
title Open foundation models for Azerbaijani language
topic Computation and Language
url https://arxiv.org/abs/2407.02337