FAIRGAMER: Evaluating Social Biases in LLM-Based Video Game NPCs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shi, Bingkang, Huang, Jen-tse, Luo, Long, Zong, Tianyu, Yi, Hongzhu, Wang, Yuanxiang, Hu, Songlin, Zhang, Xiaodan, Yao, Zhongjiang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915741899423744
author Shi, Bingkang
Huang, Jen-tse
Luo, Long
Zong, Tianyu
Yi, Hongzhu
Wang, Yuanxiang
Hu, Songlin
Zhang, Xiaodan
Yao, Zhongjiang
author_facet Shi, Bingkang
Huang, Jen-tse
Luo, Long
Zong, Tianyu
Yi, Hongzhu
Wang, Yuanxiang
Hu, Songlin
Zhang, Xiaodan
Yao, Zhongjiang
contents Large Language Models (LLMs) have increasingly enhanced or replaced traditional Non-Player Characters (NPCs) in video games. However, these LLM-based NPCs inherit underlying social biases (e.g., race or class), posing fairness risks during in-game interactions. To address the limited exploration of this issue, we introduce FairGamer, the first benchmark to evaluate social biases across three interaction patterns: transaction, cooperation, and competition. FairGamer assesses four bias types, including class, race, age, and nationality, across 12 distinct evaluation tasks using a novel metric, FairMCV. Our evaluation of seven frontier LLMs reveals that: (1) models exhibit biased decision-making, with Grok-4-Fast demonstrating the highest bias (average FairMCV = 76.9%); and (2) larger LLMs display more severe social biases, suggesting that increased model capacity inadvertently amplifies these biases. We release FairGamer at https://github.com/Anonymous999-xxx/FairGamer to facilitate future research on NPC fairness.
format Preprint
id arxiv_https___arxiv_org_abs_2508_17825
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FAIRGAMER: Evaluating Social Biases in LLM-Based Video Game NPCs
Shi, Bingkang
Huang, Jen-tse
Luo, Long
Zong, Tianyu
Yi, Hongzhu
Wang, Yuanxiang
Hu, Songlin
Zhang, Xiaodan
Yao, Zhongjiang
Artificial Intelligence
Large Language Models (LLMs) have increasingly enhanced or replaced traditional Non-Player Characters (NPCs) in video games. However, these LLM-based NPCs inherit underlying social biases (e.g., race or class), posing fairness risks during in-game interactions. To address the limited exploration of this issue, we introduce FairGamer, the first benchmark to evaluate social biases across three interaction patterns: transaction, cooperation, and competition. FairGamer assesses four bias types, including class, race, age, and nationality, across 12 distinct evaluation tasks using a novel metric, FairMCV. Our evaluation of seven frontier LLMs reveals that: (1) models exhibit biased decision-making, with Grok-4-Fast demonstrating the highest bias (average FairMCV = 76.9%); and (2) larger LLMs display more severe social biases, suggesting that increased model capacity inadvertently amplifies these biases. We release FairGamer at https://github.com/Anonymous999-xxx/FairGamer to facilitate future research on NPC fairness.
title FAIRGAMER: Evaluating Social Biases in LLM-Based Video Game NPCs
topic Artificial Intelligence
url https://arxiv.org/abs/2508.17825