Beyond Guilt: Legal Judgment Prediction with Trichotomous Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Kepu, Yang, Haoyue, Tang, Xu, Yu, Weijie, Xu, Jun
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909505212645376
author Zhang, Kepu
Yang, Haoyue
Tang, Xu
Yu, Weijie
Xu, Jun
author_facet Zhang, Kepu
Yang, Haoyue
Tang, Xu
Yu, Weijie
Xu, Jun
contents In legal practice, judges apply the trichotomous dogmatics of criminal law, sequentially assessing the elements of the offense, unlawfulness, and culpability to determine whether an individual's conduct constitutes a crime. Although current legal large language models (LLMs) show promising accuracy in judgment prediction, they lack trichotomous reasoning capabilities due to the absence of an appropriate benchmark dataset, preventing them from predicting innocent outcomes. As a result, every input is automatically assigned a charge, limiting their practical utility in legal contexts. To bridge this gap, we introduce LJPIV, the first benchmark dataset for Legal Judgment Prediction with Innocent Verdicts. Adhering to the trichotomous dogmatics, we extend three widely-used legal datasets through LLM-based augmentation and manual verification. Our experiments with state-of-the-art legal LLMs and novel strategies that integrate trichotomous reasoning into zero-shot prompting and fine-tuning reveal: (1) current legal LLMs have significant room for improvement, with even the best models achieving an F1 score of less than 0.3 on LJPIV; and (2) our strategies notably enhance both in-domain and cross-domain judgment prediction accuracy, especially for cases resulting in an innocent verdict.
format Preprint
id arxiv_https___arxiv_org_abs_2412_14588
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Beyond Guilt: Legal Judgment Prediction with Trichotomous Reasoning
Zhang, Kepu
Yang, Haoyue
Tang, Xu
Yu, Weijie
Xu, Jun
Computation and Language
In legal practice, judges apply the trichotomous dogmatics of criminal law, sequentially assessing the elements of the offense, unlawfulness, and culpability to determine whether an individual's conduct constitutes a crime. Although current legal large language models (LLMs) show promising accuracy in judgment prediction, they lack trichotomous reasoning capabilities due to the absence of an appropriate benchmark dataset, preventing them from predicting innocent outcomes. As a result, every input is automatically assigned a charge, limiting their practical utility in legal contexts. To bridge this gap, we introduce LJPIV, the first benchmark dataset for Legal Judgment Prediction with Innocent Verdicts. Adhering to the trichotomous dogmatics, we extend three widely-used legal datasets through LLM-based augmentation and manual verification. Our experiments with state-of-the-art legal LLMs and novel strategies that integrate trichotomous reasoning into zero-shot prompting and fine-tuning reveal: (1) current legal LLMs have significant room for improvement, with even the best models achieving an F1 score of less than 0.3 on LJPIV; and (2) our strategies notably enhance both in-domain and cross-domain judgment prediction accuracy, especially for cases resulting in an innocent verdict.
title Beyond Guilt: Legal Judgment Prediction with Trichotomous Reasoning
topic Computation and Language
url https://arxiv.org/abs/2412.14588