When Fairness Isn't Statistical: The Limits of Machine Learning in Evaluating Legal Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Barale, Claire, Rovatsos, Michael, Bhuta, Nehal |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LexTime: A Benchmark for Temporal Ordering of Legal Events
by: Barale, Claire, et al.
Published: (2025)
by: Barale, Claire, et al.
Published: (2025)
When Correct Isn't Usable: Improving Structured Output Reliability in Small Language Models
by: Galeone, Cosimo, et al.
Published: (2026)
by: Galeone, Cosimo, et al.
Published: (2026)
Unlearning Isn't Deletion: Investigating Reversibility of Machine Unlearning in LLMs
by: Xu, Xiaoyu, et al.
Published: (2025)
by: Xu, Xiaoyu, et al.
Published: (2025)
When Bigger Isn't Better: A Comprehensive Fairness Evaluation of Political Bias in Multi-News Summarisation
by: Huang, Nannan, et al.
Published: (2026)
by: Huang, Nannan, et al.
Published: (2026)
Inverse Scaling: When Bigger Isn't Better
by: McKenzie, Ian R., et al.
Published: (2023)
by: McKenzie, Ian R., et al.
Published: (2023)
ML Interpretability: Simple Isn't Easy
by: Räz, Tim
Published: (2022)
by: Räz, Tim
Published: (2022)
Relevance Isn't All You Need: Scaling RAG Systems With Inference-Time Compute Via Multi-Criteria Reranking
by: LeVine, Will, et al.
Published: (2025)
by: LeVine, Will, et al.
Published: (2025)
Evaluating the Limits of Large Language Models in Multilingual Legal Reasoning
by: Ioannou, Antreas, et al.
Published: (2025)
by: Ioannou, Antreas, et al.
Published: (2025)
Strong Reasoning Isn't Enough: Evaluating Evidence Elicitation in Interactive Diagnosis
by: Long, Zhuohan, et al.
Published: (2026)
by: Long, Zhuohan, et al.
Published: (2026)
What You Read Isn't What You Hear: Linguistic Sensitivity in Deepfake Speech Detection
by: Nguyen, Binh, et al.
Published: (2025)
by: Nguyen, Binh, et al.
Published: (2025)
When Meaning Isn't Literal: Exploring Idiomatic Meaning Across Languages and Modalities
by: Das, Sarmistha, et al.
Published: (2026)
by: Das, Sarmistha, et al.
Published: (2026)
Reasoning Isn't Enough: Examining Truth-Bias and Sycophancy in LLMs
by: Barkett, Emilio, et al.
Published: (2025)
by: Barkett, Emilio, et al.
Published: (2025)
Analogical Reasoning Inside Large Language Models: Concept Vectors and the Limits of Abstraction
by: Opiełka, Gustaw, et al.
Published: (2025)
by: Opiełka, Gustaw, et al.
Published: (2025)
When Privacy Isn't Synthetic: Hidden Data Leakage in Generative AI Models
by: Mustaqim, S. M., et al.
Published: (2025)
by: Mustaqim, S. M., et al.
Published: (2025)
Why Isn't Relational Learning Taking Over the World?
by: Poole, David
Published: (2025)
by: Poole, David
Published: (2025)
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning
by: Lyu, Chengqi, et al.
Published: (2025)
by: Lyu, Chengqi, et al.
Published: (2025)
Explainable AI Isn't Enough! Rethinking Algorithmic Contestability
by: Freiesleben, Timo, et al.
Published: (2026)
by: Freiesleben, Timo, et al.
Published: (2026)
Seeing Isn't Knowing: Do VLMs Know When Not to Answer Spatial Questions (and Why)?
by: Zhang, Yue, et al.
Published: (2026)
by: Zhang, Yue, et al.
Published: (2026)
Word Boundary Information Isn't Useful for Encoder Language Models
by: Gow-Smith, Edward, et al.
Published: (2024)
by: Gow-Smith, Edward, et al.
Published: (2024)
Seeing Isn't Believing: Uncovering Blind Spots in Evaluator Vision-Language Models
by: Khan, Mohammed Safi Ur Rahman, et al.
Published: (2026)
by: Khan, Mohammed Safi Ur Rahman, et al.
Published: (2026)
Assessing Dialect Fairness and Robustness of Large Language Models in Reasoning Tasks
by: Lin, Fangru, et al.
Published: (2024)
by: Lin, Fangru, et al.
Published: (2024)
Knowledge Graph-Assisted LLM Post-Training for Enhanced Legal Reasoning
by: Song, Dezhao, et al.
Published: (2026)
by: Song, Dezhao, et al.
Published: (2026)
Addressing Both Statistical and Causal Gender Fairness in NLP Models
by: Chen, Hannah, et al.
Published: (2024)
by: Chen, Hannah, et al.
Published: (2024)
Transformer See, Transformer Do: Copying as an Intermediate Step in Learning Analogical Reasoning
by: Hellwig, Philipp, et al.
Published: (2026)
by: Hellwig, Philipp, et al.
Published: (2026)
Why Code, Why Now: An Information-Theoretic Perspective on the Limits of Machine Learning
by: Zhao, Zhimin
Published: (2026)
by: Zhao, Zhimin
Published: (2026)
When Actions Teach You to Think: Reasoning-Action Synergy via Reinforcement Learning in Conversational Agents
by: Rawat, Mrinal, et al.
Published: (2025)
by: Rawat, Mrinal, et al.
Published: (2025)
NyayaMind- A Framework for Transparent Legal Reasoning and Judgment Prediction in the Indian Legal System
by: Shukla, Parjanya Aditya, et al.
Published: (2026)
by: Shukla, Parjanya Aditya, et al.
Published: (2026)
AdaptThink: Reasoning Models Can Learn When to Think
by: Zhang, Jiajie, et al.
Published: (2025)
by: Zhang, Jiajie, et al.
Published: (2025)
When Do Prompting and Prefix-Tuning Work? A Theory of Capabilities and Limitations
by: Petrov, Aleksandar, et al.
Published: (2023)
by: Petrov, Aleksandar, et al.
Published: (2023)
Weak-to-Strong Generalization beyond Accuracy: a Pilot Study in Safety, Toxicity, and Legal Reasoning
by: Ye, Ruimeng, et al.
Published: (2024)
by: Ye, Ruimeng, et al.
Published: (2024)
Reasoning Beyond Limits: Advances and Open Problems for LLMs
by: Ferrag, Mohamed Amine, et al.
Published: (2025)
by: Ferrag, Mohamed Amine, et al.
Published: (2025)
Bigger Isn't Always Memorizing: Early Stopping Overparameterized Diffusion Models
by: Favero, Alessandro, et al.
Published: (2025)
by: Favero, Alessandro, et al.
Published: (2025)
FEval-TTC: Fair Evaluation Protocol for Test-Time Compute
by: Rumiantsev, Pavel, et al.
Published: (2025)
by: Rumiantsev, Pavel, et al.
Published: (2025)
CEB: Compositional Evaluation Benchmark for Fairness in Large Language Models
by: Wang, Song, et al.
Published: (2024)
by: Wang, Song, et al.
Published: (2024)
Being Kind Isn't Always Being Safe: Diagnosing Affective Hallucination in LLMs
by: Kim, Sewon, et al.
Published: (2025)
by: Kim, Sewon, et al.
Published: (2025)
GTBench: Uncovering the Strategic Reasoning Limitations of LLMs via Game-Theoretic Evaluations
by: Duan, Jinhao, et al.
Published: (2024)
by: Duan, Jinhao, et al.
Published: (2024)
Causal Reasoning Favors Encoders: On The Limits of Decoder-Only Models
by: Roy, Amartya, et al.
Published: (2025)
by: Roy, Amartya, et al.
Published: (2025)
MMLU-Pro+: Evaluating Higher-Order Reasoning and Shortcut Learning in LLMs
by: Taghanaki, Saeid Asgari, et al.
Published: (2024)
by: Taghanaki, Saeid Asgari, et al.
Published: (2024)
Systematic Evaluation of Machine-Generated Reasoning and PHQ-9 Labeling for Depression Detection Using Large Language Models
by: Shao, Zongru, et al.
Published: (2025)
by: Shao, Zongru, et al.
Published: (2025)
IL-TUR: Benchmark for Indian Legal Text Understanding and Reasoning
by: Joshi, Abhinav, et al.
Published: (2024)
by: Joshi, Abhinav, et al.
Published: (2024)
Similar Items
-
LexTime: A Benchmark for Temporal Ordering of Legal Events
by: Barale, Claire, et al.
Published: (2025) -
When Correct Isn't Usable: Improving Structured Output Reliability in Small Language Models
by: Galeone, Cosimo, et al.
Published: (2026) -
Unlearning Isn't Deletion: Investigating Reversibility of Machine Unlearning in LLMs
by: Xu, Xiaoyu, et al.
Published: (2025) -
When Bigger Isn't Better: A Comprehensive Fairness Evaluation of Political Bias in Multi-News Summarisation
by: Huang, Nannan, et al.
Published: (2026) -
Inverse Scaling: When Bigger Isn't Better
by: McKenzie, Ian R., et al.
Published: (2023)