Towards Typologically Aware Rescoring to Mitigate Unfaithfulness in Lower-Resource Languages

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Chan, Tsan Tsai, Tong, Xin, Hoang, Thi Thu Uyen, Tepnadze, Barbare, Stempniak, Wojciech
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910864837181440
author Chan, Tsan Tsai
Tong, Xin
Hoang, Thi Thu Uyen
Tepnadze, Barbare
Stempniak, Wojciech
author_facet Chan, Tsan Tsai
Tong, Xin
Hoang, Thi Thu Uyen
Tepnadze, Barbare
Stempniak, Wojciech
contents Multilingual large language models (LLMs) are known to more frequently generate non-faithful output in resource-constrained languages (Guerreiro et al., 2023 - arXiv:2303.16104), potentially because these typologically diverse languages are underrepresented in their training data. To mitigate unfaithfulness in such settings, we propose using computationally light auxiliary models to rescore the outputs of larger architectures. As proof of the feasibility of such an approach, we show that monolingual 4-layer BERT models pretrained from scratch on less than 700 MB of data without fine-tuning are able to identify faithful summaries with a mean accuracy of 88.33% in three genetically unrelated languages that differ in their morphological complexity - Vietnamese, Polish and Georgian. The same hyperparameter combination moreover generalises well to three other tasks, suggesting applications for rescoring beyond improving faithfulness. In order to inform typologically aware model selection, we also investigate how morphological complexity interacts with regularisation, model depth and training objectives, ultimately demonstrating that morphologically complex languages are more likely to benefit from dropout, while across languages downstream performance is enhanced most by shallow architectures as well as training using the standard BERT objectives.
format Preprint
id arxiv_https___arxiv_org_abs_2502_17664
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Towards Typologically Aware Rescoring to Mitigate Unfaithfulness in Lower-Resource Languages
Chan, Tsan Tsai
Tong, Xin
Hoang, Thi Thu Uyen
Tepnadze, Barbare
Stempniak, Wojciech
Computation and Language
Multilingual large language models (LLMs) are known to more frequently generate non-faithful output in resource-constrained languages (Guerreiro et al., 2023 - arXiv:2303.16104), potentially because these typologically diverse languages are underrepresented in their training data. To mitigate unfaithfulness in such settings, we propose using computationally light auxiliary models to rescore the outputs of larger architectures. As proof of the feasibility of such an approach, we show that monolingual 4-layer BERT models pretrained from scratch on less than 700 MB of data without fine-tuning are able to identify faithful summaries with a mean accuracy of 88.33% in three genetically unrelated languages that differ in their morphological complexity - Vietnamese, Polish and Georgian. The same hyperparameter combination moreover generalises well to three other tasks, suggesting applications for rescoring beyond improving faithfulness. In order to inform typologically aware model selection, we also investigate how morphological complexity interacts with regularisation, model depth and training objectives, ultimately demonstrating that morphologically complex languages are more likely to benefit from dropout, while across languages downstream performance is enhanced most by shallow architectures as well as training using the standard BERT objectives.
title Towards Typologically Aware Rescoring to Mitigate Unfaithfulness in Lower-Resource Languages
topic Computation and Language
url https://arxiv.org/abs/2502.17664