LP-Eval: Rubric and Dataset for Measuring the Quality of Legal Proposition Generation

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Xu, Shanshan, Lindholm, Johan, Raina, Amogh, Olsen, Henrik Palmer, Hershcovich, Daniel
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866916027939422208
author Xu, Shanshan
Lindholm, Johan
Raina, Amogh
Olsen, Henrik Palmer
Hershcovich, Daniel
author_facet Xu, Shanshan
Lindholm, Johan
Raina, Amogh
Olsen, Henrik Palmer
Hershcovich, Daniel
contents Legal proposition generation is central to legal reasoning and doctrinal scholarship, yet remain under-examined in Legal NLP. This paper investigates the automatic generation and evaluation of legal propositions from decisions of the Court of Justice of the European Union using large language models (LLMs). We introduce LP-Eval, a three-step evaluation rubric co-designed with legal experts that decomposes legal proposition quality into formal validity and substantive dimensions. Using this rubric, we release a dataset of two experts' annotations for 100 LLM-generated legal propositions. Our results show that LLMs can generate predominantly well-formed and high-quality propositions, while expert evaluations reveal higher quality for propositions derived from well established cases than from recent ones. We further examine LLMs as evaluators and find that rubric-guided LLM judgments align more closely with expert assessments than direct overall scoring, but remain insensitive to finer-grained distinctions captured by human experts.
format Preprint
id arxiv_https___arxiv_org_abs_2605_19815
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle LP-Eval: Rubric and Dataset for Measuring the Quality of Legal Proposition Generation
Xu, Shanshan
Lindholm, Johan
Raina, Amogh
Olsen, Henrik Palmer
Hershcovich, Daniel
Computation and Language
Artificial Intelligence
Legal proposition generation is central to legal reasoning and doctrinal scholarship, yet remain under-examined in Legal NLP. This paper investigates the automatic generation and evaluation of legal propositions from decisions of the Court of Justice of the European Union using large language models (LLMs). We introduce LP-Eval, a three-step evaluation rubric co-designed with legal experts that decomposes legal proposition quality into formal validity and substantive dimensions. Using this rubric, we release a dataset of two experts' annotations for 100 LLM-generated legal propositions. Our results show that LLMs can generate predominantly well-formed and high-quality propositions, while expert evaluations reveal higher quality for propositions derived from well established cases than from recent ones. We further examine LLMs as evaluators and find that rubric-guided LLM judgments align more closely with expert assessments than direct overall scoring, but remain insensitive to finer-grained distinctions captured by human experts.
title LP-Eval: Rubric and Dataset for Measuring the Quality of Legal Proposition Generation
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2605.19815