Cross-lingual Transfer of Reward Models in Multilingual Alignment

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hong, Jiwoo, Lee, Noah, Martínez-Castaño, Rodrigo, Rodríguez, César, Thorne, James
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915117116948480
author Hong, Jiwoo
Lee, Noah
Martínez-Castaño, Rodrigo
Rodríguez, César
Thorne, James
author_facet Hong, Jiwoo
Lee, Noah
Martínez-Castaño, Rodrigo
Rodríguez, César
Thorne, James
contents Reinforcement learning with human feedback (RLHF) is shown to largely benefit from precise reward models (RMs). However, recent studies in reward modeling schemes are skewed towards English, limiting the applicability of RLHF in multilingual alignments. In this work, we investigate the cross-lingual transfer of RMs trained in diverse languages, primarily from English. Our experimental results demonstrate the strong cross-lingual transfer of English RMs, exceeding target language RMs by 3~4% average increase in Multilingual RewardBench. Furthermore, we analyze the cross-lingual transfer of RMs through the representation shifts. Finally, we perform multilingual alignment to exemplify how cross-lingual transfer in RM propagates to enhanced multilingual instruction-following capability, along with extensive analyses on off-the-shelf RMs. We release the code, model, and data.
format Preprint
id arxiv_https___arxiv_org_abs_2410_18027
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Cross-lingual Transfer of Reward Models in Multilingual Alignment
Hong, Jiwoo
Lee, Noah
Martínez-Castaño, Rodrigo
Rodríguez, César
Thorne, James
Computation and Language
Artificial Intelligence
Reinforcement learning with human feedback (RLHF) is shown to largely benefit from precise reward models (RMs). However, recent studies in reward modeling schemes are skewed towards English, limiting the applicability of RLHF in multilingual alignments. In this work, we investigate the cross-lingual transfer of RMs trained in diverse languages, primarily from English. Our experimental results demonstrate the strong cross-lingual transfer of English RMs, exceeding target language RMs by 3~4% average increase in Multilingual RewardBench. Furthermore, we analyze the cross-lingual transfer of RMs through the representation shifts. Finally, we perform multilingual alignment to exemplify how cross-lingual transfer in RM propagates to enhanced multilingual instruction-following capability, along with extensive analyses on off-the-shelf RMs. We release the code, model, and data.
title Cross-lingual Transfer of Reward Models in Multilingual Alignment
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2410.18027