Fluent Alignment with Disfluent Judges: Post-training for Lower-resource Languages

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Samuel, David, Øvrelid, Lilja, Velldal, Erik, Kutuzov, Andrey
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914426240630784
author Samuel, David
Øvrelid, Lilja
Velldal, Erik
Kutuzov, Andrey
author_facet Samuel, David
Øvrelid, Lilja
Velldal, Erik
Kutuzov, Andrey
contents We propose a post-training method for lower-resource languages that preserves the fluency of language models even when aligned by disfluent reward models. Preference optimization is now a well-researched topic, but previous work has mostly addressed models for English and Chinese. Lower-resource languages lack both datasets written by native speakers and instruction-tuned language models capable of generating fluent synthetic data. To address this, we focus on developing a fluent preference-aligned language model without any instruction-tuning data in the target language. Our approach uses an on-policy training method, which we compare with two common alternatives: supervised finetuning on machine-translated data and multilingual finetuning. We conduct a case study on Norwegian Bokmål and evaluate fluency through native-speaker assessments. The results show that the on-policy aspect is crucial and outperforms the alternatives without relying on any hard-to-obtain data.
format Preprint
id arxiv_https___arxiv_org_abs_2512_08777
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Fluent Alignment with Disfluent Judges: Post-training for Lower-resource Languages
Samuel, David
Øvrelid, Lilja
Velldal, Erik
Kutuzov, Andrey
Computation and Language
Artificial Intelligence
We propose a post-training method for lower-resource languages that preserves the fluency of language models even when aligned by disfluent reward models. Preference optimization is now a well-researched topic, but previous work has mostly addressed models for English and Chinese. Lower-resource languages lack both datasets written by native speakers and instruction-tuned language models capable of generating fluent synthetic data. To address this, we focus on developing a fluent preference-aligned language model without any instruction-tuning data in the target language. Our approach uses an on-policy training method, which we compare with two common alternatives: supervised finetuning on machine-translated data and multilingual finetuning. We conduct a case study on Norwegian Bokmål and evaluate fluency through native-speaker assessments. The results show that the on-policy aspect is crucial and outperforms the alternatives without relying on any hard-to-obtain data.
title Fluent Alignment with Disfluent Judges: Post-training for Lower-resource Languages
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2512.08777