LLMs on Trial: Evaluating Judicial Fairness for Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hu, Yiran, Xue, Zongyue, Li, Haitao, Zheng, Siyuan, Chen, Qingjing, Wang, Shaochun, Zhang, Xihan, Zheng, Ning, Liu, Yun, Ai, Qingyao, Liu, Yiqun, Clarke, Charles L. A., Shen, Weixing
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915422821941248
author Hu, Yiran
Xue, Zongyue
Li, Haitao
Zheng, Siyuan
Chen, Qingjing
Wang, Shaochun
Zhang, Xihan
Zheng, Ning
Liu, Yun
Ai, Qingyao
Liu, Yiqun
Clarke, Charles L. A.
Shen, Weixing
author_facet Hu, Yiran
Xue, Zongyue
Li, Haitao
Zheng, Siyuan
Chen, Qingjing
Wang, Shaochun
Zhang, Xihan
Zheng, Ning
Liu, Yun
Ai, Qingyao
Liu, Yiqun
Clarke, Charles L. A.
Shen, Weixing
contents Large Language Models (LLMs) are increasingly used in high-stakes fields where their decisions impact rights and equity. However, LLMs' judicial fairness and implications for social justice remain underexplored. When LLMs act as judges, the ability to fairly resolve judicial issues is a prerequisite to ensure their trustworthiness. Based on theories of judicial fairness, we construct a comprehensive framework to measure LLM fairness, leading to a selection of 65 labels and 161 corresponding values. Applying this framework to the judicial system, we compile an extensive dataset, JudiFair, comprising 177,100 unique case facts. To achieve robust statistical inference, we develop three evaluation metrics, inconsistency, bias, and imbalanced inaccuracy, and introduce a method to assess the overall fairness of multiple LLMs across various labels. Through experiments with 16 LLMs, we uncover pervasive inconsistency, bias, and imbalanced inaccuracy across models, underscoring severe LLM judicial unfairness. Particularly, LLMs display notably more pronounced biases on demographic labels, with slightly less bias on substance labels compared to procedure ones. Interestingly, increased inconsistency correlates with reduced biases, but more accurate predictions exacerbate biases. While we find that adjusting the temperature parameter can influence LLM fairness, model size, release date, and country of origin do not exhibit significant effects on judicial fairness. Accordingly, we introduce a publicly available toolkit containing all datasets and code, designed to support future research in evaluating and improving LLM fairness.
format Preprint
id arxiv_https___arxiv_org_abs_2507_10852
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LLMs on Trial: Evaluating Judicial Fairness for Large Language Models
Hu, Yiran
Xue, Zongyue
Li, Haitao
Zheng, Siyuan
Chen, Qingjing
Wang, Shaochun
Zhang, Xihan
Zheng, Ning
Liu, Yun
Ai, Qingyao
Liu, Yiqun
Clarke, Charles L. A.
Shen, Weixing
Computation and Language
Large Language Models (LLMs) are increasingly used in high-stakes fields where their decisions impact rights and equity. However, LLMs' judicial fairness and implications for social justice remain underexplored. When LLMs act as judges, the ability to fairly resolve judicial issues is a prerequisite to ensure their trustworthiness. Based on theories of judicial fairness, we construct a comprehensive framework to measure LLM fairness, leading to a selection of 65 labels and 161 corresponding values. Applying this framework to the judicial system, we compile an extensive dataset, JudiFair, comprising 177,100 unique case facts. To achieve robust statistical inference, we develop three evaluation metrics, inconsistency, bias, and imbalanced inaccuracy, and introduce a method to assess the overall fairness of multiple LLMs across various labels. Through experiments with 16 LLMs, we uncover pervasive inconsistency, bias, and imbalanced inaccuracy across models, underscoring severe LLM judicial unfairness. Particularly, LLMs display notably more pronounced biases on demographic labels, with slightly less bias on substance labels compared to procedure ones. Interestingly, increased inconsistency correlates with reduced biases, but more accurate predictions exacerbate biases. While we find that adjusting the temperature parameter can influence LLM fairness, model size, release date, and country of origin do not exhibit significant effects on judicial fairness. Accordingly, we introduce a publicly available toolkit containing all datasets and code, designed to support future research in evaluating and improving LLM fairness.
title LLMs on Trial: Evaluating Judicial Fairness for Large Language Models
topic Computation and Language
url https://arxiv.org/abs/2507.10852