JustEva: A Toolkit to Evaluate LLM Fairness in Legal Knowledge Inference

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Xue, Zongyue, Zheng, Siyuan, Wang, Shaochun, Hu, Yiran, Wang, Shenran, Yao, Yuxin, Li, Haitao, Ai, Qingyao, Liu, Yiqun, Liu, Yun, Shen, Weixing
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912587835244544
author Xue, Zongyue
Zheng, Siyuan
Wang, Shaochun
Hu, Yiran
Wang, Shenran
Yao, Yuxin
Li, Haitao
Ai, Qingyao
Liu, Yiqun
Liu, Yun
Shen, Weixing
author_facet Xue, Zongyue
Zheng, Siyuan
Wang, Shaochun
Hu, Yiran
Wang, Shenran
Yao, Yuxin
Li, Haitao
Ai, Qingyao
Liu, Yiqun
Liu, Yun
Shen, Weixing
contents The integration of Large Language Models (LLMs) into legal practice raises pressing concerns about judicial fairness, particularly due to the nature of their "black-box" processes. This study introduces JustEva, a comprehensive, open-source evaluation toolkit designed to measure LLM fairness in legal tasks. JustEva features several advantages: (1) a structured label system covering 65 extra-legal factors; (2) three core fairness metrics - inconsistency, bias, and imbalanced inaccuracy; (3) robust statistical inference methods; and (4) informative visualizations. The toolkit supports two types of experiments, enabling a complete evaluation workflow: (1) generating structured outputs from LLMs using a provided dataset, and (2) conducting statistical analysis and inference on LLMs' outputs through regression and other statistical methods. Empirical application of JustEva reveals significant fairness deficiencies in current LLMs, highlighting the lack of fair and trustworthy LLM legal tools. JustEva offers a convenient tool and methodological foundation for evaluating and improving algorithmic fairness in the legal domain.
format Preprint
id arxiv_https___arxiv_org_abs_2509_12104
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle JustEva: A Toolkit to Evaluate LLM Fairness in Legal Knowledge Inference
Xue, Zongyue
Zheng, Siyuan
Wang, Shaochun
Hu, Yiran
Wang, Shenran
Yao, Yuxin
Li, Haitao
Ai, Qingyao
Liu, Yiqun
Liu, Yun
Shen, Weixing
Artificial Intelligence
The integration of Large Language Models (LLMs) into legal practice raises pressing concerns about judicial fairness, particularly due to the nature of their "black-box" processes. This study introduces JustEva, a comprehensive, open-source evaluation toolkit designed to measure LLM fairness in legal tasks. JustEva features several advantages: (1) a structured label system covering 65 extra-legal factors; (2) three core fairness metrics - inconsistency, bias, and imbalanced inaccuracy; (3) robust statistical inference methods; and (4) informative visualizations. The toolkit supports two types of experiments, enabling a complete evaluation workflow: (1) generating structured outputs from LLMs using a provided dataset, and (2) conducting statistical analysis and inference on LLMs' outputs through regression and other statistical methods. Empirical application of JustEva reveals significant fairness deficiencies in current LLMs, highlighting the lack of fair and trustworthy LLM legal tools. JustEva offers a convenient tool and methodological foundation for evaluating and improving algorithmic fairness in the legal domain.
title JustEva: A Toolkit to Evaluate LLM Fairness in Legal Knowledge Inference
topic Artificial Intelligence
url https://arxiv.org/abs/2509.12104