Exploring Pretraining via Active Forgetting for Improving Cross Lingual Transfer for Decoder Language Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Aggarwal, Divyanshu, Sathe, Ashutosh, Sitaram, Sunayana
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913850486423552
author Aggarwal, Divyanshu
Sathe, Ashutosh
Sitaram, Sunayana
author_facet Aggarwal, Divyanshu
Sathe, Ashutosh
Sitaram, Sunayana
contents Large Language Models (LLMs) demonstrate exceptional capabilities in a multitude of NLP tasks. However, the efficacy of such models to languages other than English is often limited. Prior works have shown that encoder-only models such as BERT or XLM-RoBERTa show impressive cross lingual transfer of their capabilities from English to other languages. In this work, we propose a pretraining strategy that uses active forgetting to achieve similar cross lingual transfer in decoder-only LLMs. We show that LLMs pretrained with active forgetting are highly effective when adapting to new and unseen languages. Through extensive experimentation, we find that LLMs pretrained with active forgetting are able to learn better multilingual representations which translates to better performance in many downstream tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2410_16168
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Exploring Pretraining via Active Forgetting for Improving Cross Lingual Transfer for Decoder Language Models
Aggarwal, Divyanshu
Sathe, Ashutosh
Sitaram, Sunayana
Computation and Language
Large Language Models (LLMs) demonstrate exceptional capabilities in a multitude of NLP tasks. However, the efficacy of such models to languages other than English is often limited. Prior works have shown that encoder-only models such as BERT or XLM-RoBERTa show impressive cross lingual transfer of their capabilities from English to other languages. In this work, we propose a pretraining strategy that uses active forgetting to achieve similar cross lingual transfer in decoder-only LLMs. We show that LLMs pretrained with active forgetting are highly effective when adapting to new and unseen languages. Through extensive experimentation, we find that LLMs pretrained with active forgetting are able to learn better multilingual representations which translates to better performance in many downstream tasks.
title Exploring Pretraining via Active Forgetting for Improving Cross Lingual Transfer for Decoder Language Models
topic Computation and Language
url https://arxiv.org/abs/2410.16168