LLMs on Trial: Evaluating Judicial Fairness for Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Hu, Yiran, Xue, Zongyue, Li, Haitao, Zheng, Siyuan, Chen, Qingjing, Wang, Shaochun, Zhang, Xihan, Zheng, Ning, Liu, Yun, Ai, Qingyao, Liu, Yiqun, Clarke, Charles L. A., Shen, Weixing |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
JustEva: A Toolkit to Evaluate LLM Fairness in Legal Knowledge Inference
by: Xue, Zongyue, et al.
Published: (2025)
by: Xue, Zongyue, et al.
Published: (2025)
J&H: Evaluating the Robustness of Large Language Models Under Knowledge-Injection Attacks in Legal Domain
by: Hu, Yiran, et al.
Published: (2025)
by: Hu, Yiran, et al.
Published: (2025)
STARD: A Chinese Statute Retrieval Dataset with Real Queries Issued by Non-professionals
by: Su, Weihang, et al.
Published: (2024)
by: Su, Weihang, et al.
Published: (2024)
Evaluation of Large Language Models in Legal Applications: Challenges, Methods, and Future Directions
by: Hu, Yiran, et al.
Published: (2026)
by: Hu, Yiran, et al.
Published: (2026)
Overview of the NTCIR-18 Automatic Evaluation of LLMs (AEOLLM) Task
by: Chen, Junjie, et al.
Published: (2025)
by: Chen, Junjie, et al.
Published: (2025)
PRE: A Peer Review Based Large Language Model Evaluator
by: Chu, Zhumin, et al.
Published: (2024)
by: Chu, Zhumin, et al.
Published: (2024)
ATACompressor: Adaptive Task-Aware Compression for Efficient Long-Context Processing in LLMs
by: Li, Xuancheng, et al.
Published: (2026)
by: Li, Xuancheng, et al.
Published: (2026)
Beyond Experience Retrieval: Learning to Generate Utility-Optimized Structured Experience for Frozen LLMs
by: Li, Xuancheng, et al.
Published: (2026)
by: Li, Xuancheng, et al.
Published: (2026)
Evaluation Ethics of LLMs in Legal Domain
by: Zhang, Ruizhe, et al.
Published: (2024)
by: Zhang, Ruizhe, et al.
Published: (2024)
LexEval: A Comprehensive Chinese Legal Benchmark for Evaluating Large Language Models
by: Li, Haitao, et al.
Published: (2024)
by: Li, Haitao, et al.
Published: (2024)
Joint Evaluation of Answer and Reasoning Consistency for Hallucination Detection in Large Reasoning Models
by: Wang, Changyue, et al.
Published: (2025)
by: Wang, Changyue, et al.
Published: (2025)
TEC: A Collection of Human Trial-and-error Trajectories for Problem Solving
by: Zhang, Xinkai, et al.
Published: (2026)
by: Zhang, Xinkai, et al.
Published: (2026)
JUREX-4E: Juridical Expert-Annotated Four-Element Knowledge Base for Legal Reasoning
by: Liu, Huanghai, et al.
Published: (2025)
by: Liu, Huanghai, et al.
Published: (2025)
LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods
by: Li, Haitao, et al.
Published: (2024)
by: Li, Haitao, et al.
Published: (2024)
Equity vs. Equality: Optimizing Ranking Fairness for Tailored Provider Needs
by: Tu, Yiteng, et al.
Published: (2026)
by: Tu, Yiteng, et al.
Published: (2026)
Enhancing LLM-Based Agents via Global Planning and Hierarchical Execution
by: Chen, Junjie, et al.
Published: (2025)
by: Chen, Junjie, et al.
Published: (2025)
MulFeRL: Enhancing Reinforcement Learning with Verbal Feedback in a Multi-turn Loop
by: Li, Xuancheng, et al.
Published: (2026)
by: Li, Xuancheng, et al.
Published: (2026)
Beyond Exposure: Optimizing Ranking Fairness with Non-linear Time-Income Functions
by: Li, Xuancheng, et al.
Published: (2026)
by: Li, Xuancheng, et al.
Published: (2026)
Improve Large Language Model Systems with User Logs
by: Wang, Changyue, et al.
Published: (2026)
by: Wang, Changyue, et al.
Published: (2026)
Foundations of GenIR
by: Ai, Qingyao, et al.
Published: (2025)
by: Ai, Qingyao, et al.
Published: (2025)
CalibraEval: Calibrating Prediction Distribution to Mitigate Selection Bias in LLMs-as-Judges
by: Li, Haitao, et al.
Published: (2024)
by: Li, Haitao, et al.
Published: (2024)
Unsupervised Real-Time Hallucination Detection based on the Internal States of Large Language Models
by: Su, Weihang, et al.
Published: (2024)
by: Su, Weihang, et al.
Published: (2024)
Option-ID Based Elimination For Multiple Choice Questions
by: Zhu, Zhenhao, et al.
Published: (2025)
by: Zhu, Zhenhao, et al.
Published: (2025)
DRAGIN: Dynamic Retrieval Augmented Generation based on the Information Needs of Large Language Models
by: Su, Weihang, et al.
Published: (2024)
by: Su, Weihang, et al.
Published: (2024)
Multi-Field Tool Retrieval
by: Tang, Yichen, et al.
Published: (2026)
by: Tang, Yichen, et al.
Published: (2026)
CrossPT-EEG: A Benchmark for Cross-Participant and Cross-Time Generalization of EEG-based Visual Decoding
by: Zhu, Shuqi, et al.
Published: (2024)
by: Zhu, Shuqi, et al.
Published: (2024)
Decoupling Knowledge and Task Subspaces for Composable Parametric Retrieval Augmented Generation
by: Su, Weihang, et al.
Published: (2026)
by: Su, Weihang, et al.
Published: (2026)
Evaluating Intelligence via Trial and Error
by: Zhan, Jingtao, et al.
Published: (2025)
by: Zhan, Jingtao, et al.
Published: (2025)
BLADE: Enhancing Black-box Large Language Models with Small Domain-Specific Models
by: Li, Haitao, et al.
Published: (2024)
by: Li, Haitao, et al.
Published: (2024)
SelfRACG: Enabling LLMs to Self-Express and Retrieve for Code Generation
by: Dong, Qian, et al.
Published: (2025)
by: Dong, Qian, et al.
Published: (2025)
Unsupervised Large Language Model Alignment for Information Retrieval via Contrastive Feedback
by: Dong, Qian, et al.
Published: (2023)
by: Dong, Qian, et al.
Published: (2023)
LexRAG: Benchmarking Retrieval-Augmented Generation in Multi-Turn Legal Consultation Conversation
by: Li, Haitao, et al.
Published: (2025)
by: Li, Haitao, et al.
Published: (2025)
Mitigating Entity-Level Hallucination in Large Language Models
by: Su, Weihang, et al.
Published: (2024)
by: Su, Weihang, et al.
Published: (2024)
Towards an In-Depth Comprehension of Case Relevance for Better Legal Retrieval
by: Li, Haitao, et al.
Published: (2024)
by: Li, Haitao, et al.
Published: (2024)
LegalAgentBench: Evaluating LLM Agents in Legal Domain
by: Li, Haitao, et al.
Published: (2024)
by: Li, Haitao, et al.
Published: (2024)
I^3 Retriever: Incorporating Implicit Interaction in Pre-trained Language Models for Passage Retrieval
by: Dong, Qian, et al.
Published: (2023)
by: Dong, Qian, et al.
Published: (2023)
Legal Rule Induction: Towards Generalizable Principle Discovery from Analogous Judicial Precedents
by: Fan, Wei, et al.
Published: (2025)
by: Fan, Wei, et al.
Published: (2025)
Parametric Social Identity Injection and Diversification in Public Opinion Simulation
by: Wang, Hexi, et al.
Published: (2026)
by: Wang, Hexi, et al.
Published: (2026)
RbFT: Robust Fine-tuning for Retrieval-Augmented Generation against Retrieval Defects
by: Tu, Yiteng, et al.
Published: (2025)
by: Tu, Yiteng, et al.
Published: (2025)
Knowledge Editing through Chain-of-Thought
by: Wang, Changyue, et al.
Published: (2024)
by: Wang, Changyue, et al.
Published: (2024)
Similar Items
-
JustEva: A Toolkit to Evaluate LLM Fairness in Legal Knowledge Inference
by: Xue, Zongyue, et al.
Published: (2025) -
J&H: Evaluating the Robustness of Large Language Models Under Knowledge-Injection Attacks in Legal Domain
by: Hu, Yiran, et al.
Published: (2025) -
STARD: A Chinese Statute Retrieval Dataset with Real Queries Issued by Non-professionals
by: Su, Weihang, et al.
Published: (2024) -
Evaluation of Large Language Models in Legal Applications: Challenges, Methods, and Future Directions
by: Hu, Yiran, et al.
Published: (2026) -
Overview of the NTCIR-18 Automatic Evaluation of LLMs (AEOLLM) Task
by: Chen, Junjie, et al.
Published: (2025)