Fluent Alignment with Disfluent Judges: Post-training for Lower-resource Languages
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866914426240630784 |
|---|---|
| author | Samuel, David Øvrelid, Lilja Velldal, Erik Kutuzov, Andrey |
| author_facet | Samuel, David Øvrelid, Lilja Velldal, Erik Kutuzov, Andrey |
| contents | We propose a post-training method for lower-resource languages that preserves the fluency of language models even when aligned by disfluent reward models. Preference optimization is now a well-researched topic, but previous work has mostly addressed models for English and Chinese. Lower-resource languages lack both datasets written by native speakers and instruction-tuned language models capable of generating fluent synthetic data. To address this, we focus on developing a fluent preference-aligned language model without any instruction-tuning data in the target language. Our approach uses an on-policy training method, which we compare with two common alternatives: supervised finetuning on machine-translated data and multilingual finetuning. We conduct a case study on Norwegian Bokmål and evaluate fluency through native-speaker assessments. The results show that the on-policy aspect is crucial and outperforms the alternatives without relying on any hard-to-obtain data. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2512_08777 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Fluent Alignment with Disfluent Judges: Post-training for Lower-resource Languages Samuel, David Øvrelid, Lilja Velldal, Erik Kutuzov, Andrey Computation and Language Artificial Intelligence We propose a post-training method for lower-resource languages that preserves the fluency of language models even when aligned by disfluent reward models. Preference optimization is now a well-researched topic, but previous work has mostly addressed models for English and Chinese. Lower-resource languages lack both datasets written by native speakers and instruction-tuned language models capable of generating fluent synthetic data. To address this, we focus on developing a fluent preference-aligned language model without any instruction-tuning data in the target language. Our approach uses an on-policy training method, which we compare with two common alternatives: supervised finetuning on machine-translated data and multilingual finetuning. We conduct a case study on Norwegian Bokmål and evaluate fluency through native-speaker assessments. The results show that the on-policy aspect is crucial and outperforms the alternatives without relying on any hard-to-obtain data. |
| title | Fluent Alignment with Disfluent Judges: Post-training for Lower-resource Languages |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2512.08777 |