Investigating the Multilingual Calibration Effects of Language Model Instruction-Tuning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Huang, Jerry, Lu, Peng, Zeng, Qiuhao, Iwasawa, Yusuke, Matsuo, Yutaka, Chandar, Sarath, Marrese-Taylor, Edison, Li, Irene
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917183127289856
author Huang, Jerry
Lu, Peng
Zeng, Qiuhao
Iwasawa, Yusuke
Matsuo, Yutaka
Chandar, Sarath
Marrese-Taylor, Edison
Li, Irene
author_facet Huang, Jerry
Lu, Peng
Zeng, Qiuhao
Iwasawa, Yusuke
Matsuo, Yutaka
Chandar, Sarath
Marrese-Taylor, Edison
Li, Irene
contents Ensuring that deep learning models are well-calibrated in terms of their predictive uncertainty is essential in maintaining their trustworthiness and reliability, yet despite increasing advances in foundation model research, the relationship between such large language models (LLMs) and their calibration remains an open area of research. In this work, we look at a critical gap in the calibration of LLMs within multilingual settings, in an attempt to better understand how the data scarcity can potentially lead to different calibration effects and how commonly used techniques can apply in these settings. Our analysis on two multilingual benchmarks, over 29 and 42 languages respectively, reveals that even in low-resource languages, model confidence can increase significantly after instruction-tuning on high-resource language SFT datasets. However, improvements in accuracy are marginal or non-existent, resulting in mis-calibration, highlighting a critical shortcoming of standard SFT for multilingual languages. Furthermore, we observe that the use of label smoothing to be a reasonable method alleviate this concern, again without any need for low-resource SFT data, maintaining better calibration across all languages. Overall, this highlights the importance of multilingual considerations for both training and tuning LLMs in order to improve their reliability and fairness in downstream use.
format Preprint
id arxiv_https___arxiv_org_abs_2601_01362
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Investigating the Multilingual Calibration Effects of Language Model Instruction-Tuning
Huang, Jerry
Lu, Peng
Zeng, Qiuhao
Iwasawa, Yusuke
Matsuo, Yutaka
Chandar, Sarath
Marrese-Taylor, Edison
Li, Irene
Computation and Language
Machine Learning
Ensuring that deep learning models are well-calibrated in terms of their predictive uncertainty is essential in maintaining their trustworthiness and reliability, yet despite increasing advances in foundation model research, the relationship between such large language models (LLMs) and their calibration remains an open area of research. In this work, we look at a critical gap in the calibration of LLMs within multilingual settings, in an attempt to better understand how the data scarcity can potentially lead to different calibration effects and how commonly used techniques can apply in these settings. Our analysis on two multilingual benchmarks, over 29 and 42 languages respectively, reveals that even in low-resource languages, model confidence can increase significantly after instruction-tuning on high-resource language SFT datasets. However, improvements in accuracy are marginal or non-existent, resulting in mis-calibration, highlighting a critical shortcoming of standard SFT for multilingual languages. Furthermore, we observe that the use of label smoothing to be a reasonable method alleviate this concern, again without any need for low-resource SFT data, maintaining better calibration across all languages. Overall, this highlights the importance of multilingual considerations for both training and tuning LLMs in order to improve their reliability and fairness in downstream use.
title Investigating the Multilingual Calibration Effects of Language Model Instruction-Tuning
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2601.01362