CroCo: Cross-Lingual Contrastive Preference Tuning on Self-Generations

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhang, Mike, Basirat, Ali, Elliott, Desmond
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917533274079232
author Zhang, Mike
Basirat, Ali
Elliott, Desmond
author_facet Zhang, Mike
Basirat, Ali
Elliott, Desmond
contents Prior work establishes that controlled contrastiveness between self-generated responses from large language models, set via reward scores, improves downstream preference tuning in English. We extend this method to multiple languages and evaluate two models across a total of 14 high and low-resource languages on a diverse set of tasks. Our central finding is that cross-lingual contrastive preference tuning on self-generations (CroCo) transfers without language-specific preference annotation. A reward model trained on English preferences (atop a multilingual base) produces useful within-language rankings across most languages, and pairing in either a monolingual or multilingual setting improves over each model on the majority of setups while preventing the catastrophic forgetting of supervised fine-tuning. We observe that the gains require on-policy data. Off-policy responses reduce the benefit and online preference optimization fails to improve over the offline variant. Specifically, on structured tasks, our method matches or exceeds the base in 6/7 languages for EuroLLM-9B and 4/7 settings for Aya-3B. On open-ended generation, both tuned models win against their respective base across 11 evaluated languages. Overall, we show promising directions for multilingual preference tuning.
format Preprint
id arxiv_https___arxiv_org_abs_2605_26293
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle CroCo: Cross-Lingual Contrastive Preference Tuning on Self-Generations
Zhang, Mike
Basirat, Ali
Elliott, Desmond
Computation and Language
Artificial Intelligence
Prior work establishes that controlled contrastiveness between self-generated responses from large language models, set via reward scores, improves downstream preference tuning in English. We extend this method to multiple languages and evaluate two models across a total of 14 high and low-resource languages on a diverse set of tasks. Our central finding is that cross-lingual contrastive preference tuning on self-generations (CroCo) transfers without language-specific preference annotation. A reward model trained on English preferences (atop a multilingual base) produces useful within-language rankings across most languages, and pairing in either a monolingual or multilingual setting improves over each model on the majority of setups while preventing the catastrophic forgetting of supervised fine-tuning. We observe that the gains require on-policy data. Off-policy responses reduce the benefit and online preference optimization fails to improve over the offline variant. Specifically, on structured tasks, our method matches or exceeds the base in 6/7 languages for EuroLLM-9B and 4/7 settings for Aya-3B. On open-ended generation, both tuned models win against their respective base across 11 evaluated languages. Overall, we show promising directions for multilingual preference tuning.
title CroCo: Cross-Lingual Contrastive Preference Tuning on Self-Generations
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2605.26293