Evaluating Legal Reasoning Traces with Legal Issue Tree Rubrics

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lee, Jinu, On, Kyoung-Woon, Han, Simeng, Cohan, Arman, Hockenmaier, Julia
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910182178553856
author Lee, Jinu
On, Kyoung-Woon
Han, Simeng
Cohan, Arman
Hockenmaier, Julia
author_facet Lee, Jinu
On, Kyoung-Woon
Han, Simeng
Cohan, Arman
Hockenmaier, Julia
contents Evaluating the quality of LLM-generated reasoning traces in expert domains (e.g., law) is essential for ensuring credibility and explainability, yet remains challenging due to the inherent complexity of such reasoning tasks. We introduce LEGIT (LEGal Issue Trees), a novel large-scale (24K instances) expert-level legal reasoning dataset with an emphasis on reasoning trace evaluation. We convert court judgments into hierarchical trees of opposing parties' arguments and the court's conclusions, which serve as rubrics for evaluating the issue coverage and correctness of the reasoning traces. We verify the reliability of these rubrics via human expert annotations and comparison with coarse, less informative rubrics. Using the LEGIT dataset, we show that (1) LLMs' legal reasoning ability is seriously affected by both legal issue coverage and correctness, and that (2) retrieval-augmented generation (RAG) and RL with rubrics bring complementary benefits for legal reasoning abilities, where RAG improves overall reasoning capability, whereas RL improves correctness albeit with reduced coverage.
format Preprint
id arxiv_https___arxiv_org_abs_2512_01020
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Evaluating Legal Reasoning Traces with Legal Issue Tree Rubrics
Lee, Jinu
On, Kyoung-Woon
Han, Simeng
Cohan, Arman
Hockenmaier, Julia
Artificial Intelligence
Computation and Language
Evaluating the quality of LLM-generated reasoning traces in expert domains (e.g., law) is essential for ensuring credibility and explainability, yet remains challenging due to the inherent complexity of such reasoning tasks. We introduce LEGIT (LEGal Issue Trees), a novel large-scale (24K instances) expert-level legal reasoning dataset with an emphasis on reasoning trace evaluation. We convert court judgments into hierarchical trees of opposing parties' arguments and the court's conclusions, which serve as rubrics for evaluating the issue coverage and correctness of the reasoning traces. We verify the reliability of these rubrics via human expert annotations and comparison with coarse, less informative rubrics. Using the LEGIT dataset, we show that (1) LLMs' legal reasoning ability is seriously affected by both legal issue coverage and correctness, and that (2) retrieval-augmented generation (RAG) and RL with rubrics bring complementary benefits for legal reasoning abilities, where RAG improves overall reasoning capability, whereas RL improves correctness albeit with reduced coverage.
title Evaluating Legal Reasoning Traces with Legal Issue Tree Rubrics
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2512.01020