BgGPT 1.0: Extending English-centric LLMs to other languages

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Alexandrov, Anton, Raychev, Veselin, Dimitrov, Dimitar I., Zhang, Ce, Vechev, Martin, Toutanova, Kristina
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913611706793984
author Alexandrov, Anton
Raychev, Veselin
Dimitrov, Dimitar I.
Zhang, Ce
Vechev, Martin
Toutanova, Kristina
author_facet Alexandrov, Anton
Raychev, Veselin
Dimitrov, Dimitar I.
Zhang, Ce
Vechev, Martin
Toutanova, Kristina
contents We present BgGPT-Gemma-2-27B-Instruct and BgGPT-Gemma-2-9B-Instruct: continually pretrained and fine-tuned versions of Google's Gemma-2 models, specifically optimized for Bulgarian language understanding and generation. Leveraging Gemma-2's multilingual capabilities and over 100 billion tokens of Bulgarian and English text data, our models demonstrate strong performance in Bulgarian language tasks, setting a new standard for language-specific AI models. Our approach maintains the robust capabilities of the original Gemma-2 models, ensuring that the English language performance remains intact. To preserve the base model capabilities, we incorporate continual learning strategies based on recent Branch-and-Merge techniques as well as thorough curation and selection of training data. We provide detailed insights into our methodology, including the release of model weights with a commercial-friendly license, enabling broader adoption by researchers, companies, and hobbyists. Further, we establish a comprehensive set of benchmarks based on non-public educational data sources to evaluate models on Bulgarian language tasks as well as safety and chat capabilities. Our findings demonstrate the effectiveness of fine-tuning state-of-the-art models like Gemma 2 to enhance language-specific AI applications while maintaining cross-lingual capabilities.
format Preprint
id arxiv_https___arxiv_org_abs_2412_10893
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle BgGPT 1.0: Extending English-centric LLMs to other languages
Alexandrov, Anton
Raychev, Veselin
Dimitrov, Dimitar I.
Zhang, Ce
Vechev, Martin
Toutanova, Kristina
Computation and Language
Artificial Intelligence
Machine Learning
We present BgGPT-Gemma-2-27B-Instruct and BgGPT-Gemma-2-9B-Instruct: continually pretrained and fine-tuned versions of Google's Gemma-2 models, specifically optimized for Bulgarian language understanding and generation. Leveraging Gemma-2's multilingual capabilities and over 100 billion tokens of Bulgarian and English text data, our models demonstrate strong performance in Bulgarian language tasks, setting a new standard for language-specific AI models. Our approach maintains the robust capabilities of the original Gemma-2 models, ensuring that the English language performance remains intact. To preserve the base model capabilities, we incorporate continual learning strategies based on recent Branch-and-Merge techniques as well as thorough curation and selection of training data. We provide detailed insights into our methodology, including the release of model weights with a commercial-friendly license, enabling broader adoption by researchers, companies, and hobbyists. Further, we establish a comprehensive set of benchmarks based on non-public educational data sources to evaluate models on Bulgarian language tasks as well as safety and chat capabilities. Our findings demonstrate the effectiveness of fine-tuning state-of-the-art models like Gemma 2 to enhance language-specific AI applications while maintaining cross-lingual capabilities.
title BgGPT 1.0: Extending English-centric LLMs to other languages
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2412.10893