Saved in:
Bibliographic Details
Main Authors: Rojas, Mateo Alejandro, Carranza, Rafael
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2412.08955
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915060978286592
author Rojas, Mateo Alejandro
Carranza, Rafael
author_facet Rojas, Mateo Alejandro
Carranza, Rafael
contents Cross-lingual in-context learning (XICL) has emerged as a transformative paradigm for leveraging large language models (LLMs) to tackle multilingual tasks, especially for low-resource languages. However, existing approaches often rely on external retrievers or task-specific fine-tuning, limiting their scalability and generalizability. In this paper, we propose a novel self-supervised framework that harnesses the generative capabilities of LLMs to internally select and utilize task-relevant examples. Our method introduces two key objectives: a retrieval-generation alignment loss to optimize the quality of selected examples and a semantic coherence loss to ensure cross-lingual consistency. Through extensive experiments on multilingual benchmarks, our approach achieves state-of-the-art performance, significantly outperforming existing baselines. Further analysis highlights its robustness across diverse language families and its ability to generalize to unseen tasks. Human evaluations confirm the superior fluency, relevance, and semantic correctness of outputs generated by our method. This work provides a scalable, effective, and generalizable solution for cross-lingual in-context learning.
format Preprint
id arxiv_https___arxiv_org_abs_2412_08955
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Align, Generate, Learn: A Novel Closed-Loop Framework for Cross-Lingual In-Context Learning
Rojas, Mateo Alejandro
Carranza, Rafael
Computation and Language
Cross-lingual in-context learning (XICL) has emerged as a transformative paradigm for leveraging large language models (LLMs) to tackle multilingual tasks, especially for low-resource languages. However, existing approaches often rely on external retrievers or task-specific fine-tuning, limiting their scalability and generalizability. In this paper, we propose a novel self-supervised framework that harnesses the generative capabilities of LLMs to internally select and utilize task-relevant examples. Our method introduces two key objectives: a retrieval-generation alignment loss to optimize the quality of selected examples and a semantic coherence loss to ensure cross-lingual consistency. Through extensive experiments on multilingual benchmarks, our approach achieves state-of-the-art performance, significantly outperforming existing baselines. Further analysis highlights its robustness across diverse language families and its ability to generalize to unseen tasks. Human evaluations confirm the superior fluency, relevance, and semantic correctness of outputs generated by our method. This work provides a scalable, effective, and generalizable solution for cross-lingual in-context learning.
title Align, Generate, Learn: A Novel Closed-Loop Framework for Cross-Lingual In-Context Learning
topic Computation and Language
url https://arxiv.org/abs/2412.08955