Bias Mitigation or Cultural Commonsense? Evaluating LLMs with a Japanese Dataset

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yamamoto, Taisei, Kumon, Ryoma, Bollegala, Danushka, Yanaka, Hitomi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912614904233984
author Yamamoto, Taisei
Kumon, Ryoma
Bollegala, Danushka
Yanaka, Hitomi
author_facet Yamamoto, Taisei
Kumon, Ryoma
Bollegala, Danushka
Yanaka, Hitomi
contents Large language models (LLMs) exhibit social biases, prompting the development of various debiasing methods. However, debiasing methods may degrade the capabilities of LLMs. Previous research has evaluated the impact of bias mitigation primarily through tasks measuring general language understanding, which are often unrelated to social biases. In contrast, cultural commonsense is closely related to social biases, as both are rooted in social norms and values. The impact of bias mitigation on cultural commonsense in LLMs has not been well investigated. Considering this gap, we propose SOBACO (SOcial BiAs and Cultural cOmmonsense benchmark), a Japanese benchmark designed to evaluate social biases and cultural commonsense in LLMs in a unified format. We evaluate several LLMs on SOBACO to examine how debiasing methods affect cultural commonsense in LLMs. Our results reveal that the debiasing methods degrade the performance of the LLMs on the cultural commonsense task (up to 75% accuracy deterioration). These results highlight the importance of developing debiasing methods that consider the trade-off with cultural commonsense to improve fairness and utility of LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2509_24468
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Bias Mitigation or Cultural Commonsense? Evaluating LLMs with a Japanese Dataset
Yamamoto, Taisei
Kumon, Ryoma
Bollegala, Danushka
Yanaka, Hitomi
Computation and Language
Large language models (LLMs) exhibit social biases, prompting the development of various debiasing methods. However, debiasing methods may degrade the capabilities of LLMs. Previous research has evaluated the impact of bias mitigation primarily through tasks measuring general language understanding, which are often unrelated to social biases. In contrast, cultural commonsense is closely related to social biases, as both are rooted in social norms and values. The impact of bias mitigation on cultural commonsense in LLMs has not been well investigated. Considering this gap, we propose SOBACO (SOcial BiAs and Cultural cOmmonsense benchmark), a Japanese benchmark designed to evaluate social biases and cultural commonsense in LLMs in a unified format. We evaluate several LLMs on SOBACO to examine how debiasing methods affect cultural commonsense in LLMs. Our results reveal that the debiasing methods degrade the performance of the LLMs on the cultural commonsense task (up to 75% accuracy deterioration). These results highlight the importance of developing debiasing methods that consider the trade-off with cultural commonsense to improve fairness and utility of LLMs.
title Bias Mitigation or Cultural Commonsense? Evaluating LLMs with a Japanese Dataset
topic Computation and Language
url https://arxiv.org/abs/2509.24468