Exploring the Effects of Alignment on Numerical Bias in Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sato, Ayako, Kim, Hwichan, Chen, Zhousi, Mita, Masato, Komachi, Mamoru
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912848369680384
author Sato, Ayako
Kim, Hwichan
Chen, Zhousi
Mita, Masato
Komachi, Mamoru
author_facet Sato, Ayako
Kim, Hwichan
Chen, Zhousi
Mita, Masato
Komachi, Mamoru
contents "LLM-as-a-judge," which utilizes large language models (LLMs) as evaluators, has proven effective in many evaluation tasks. However, evaluator LLMs exhibit numerical bias, a phenomenon where certain evaluation scores are generated disproportionately often, leading reduced evaluation performance. This study investigates the cause of this bias. Given that most evaluator LLMs are aligned through instruction tuning and preference tuning, and that prior research suggests alignment reduces output diversity, we hypothesize that numerical bias arises from alignment. To test this, we compare outputs from pre- and post-alignment LLMs, and observe that alignment indeed increases numerical bias. We also explore mitigation strategies for post-alignment LLMs, including temperature scaling, distribution calibration, and score range adjustment. Among these, score range adjustment is most effective in reducing bias and improving performance, though still heuristic. Our findings highlight the need for further work on optimal score range selection and more robust mitigation strategies.
format Preprint
id arxiv_https___arxiv_org_abs_2601_16444
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Exploring the Effects of Alignment on Numerical Bias in Large Language Models
Sato, Ayako
Kim, Hwichan
Chen, Zhousi
Mita, Masato
Komachi, Mamoru
Computation and Language
"LLM-as-a-judge," which utilizes large language models (LLMs) as evaluators, has proven effective in many evaluation tasks. However, evaluator LLMs exhibit numerical bias, a phenomenon where certain evaluation scores are generated disproportionately often, leading reduced evaluation performance. This study investigates the cause of this bias. Given that most evaluator LLMs are aligned through instruction tuning and preference tuning, and that prior research suggests alignment reduces output diversity, we hypothesize that numerical bias arises from alignment. To test this, we compare outputs from pre- and post-alignment LLMs, and observe that alignment indeed increases numerical bias. We also explore mitigation strategies for post-alignment LLMs, including temperature scaling, distribution calibration, and score range adjustment. Among these, score range adjustment is most effective in reducing bias and improving performance, though still heuristic. Our findings highlight the need for further work on optimal score range selection and more robust mitigation strategies.
title Exploring the Effects of Alignment on Numerical Bias in Large Language Models
topic Computation and Language
url https://arxiv.org/abs/2601.16444