IL-PCSR: Legal Corpus for Prior Case and Statute Retrieval

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Paul, Shounak, Ghumare, Dhananjay, Goyal, Pawan, Ghosh, Saptarshi, Modi, Ashutosh
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914127823241216
author Paul, Shounak
Ghumare, Dhananjay
Goyal, Pawan
Ghosh, Saptarshi
Modi, Ashutosh
author_facet Paul, Shounak
Ghumare, Dhananjay
Goyal, Pawan
Ghosh, Saptarshi
Modi, Ashutosh
contents Identifying/retrieving relevant statutes and prior cases/precedents for a given legal situation are common tasks exercised by law practitioners. Researchers to date have addressed the two tasks independently, thus developing completely different datasets and models for each task; however, both retrieval tasks are inherently related, e.g., similar cases tend to cite similar statutes (due to similar factual situation). In this paper, we address this gap. We propose IL-PCR (Indian Legal corpus for Prior Case and Statute Retrieval), which is a unique corpus that provides a common testbed for developing models for both the tasks (Statute Retrieval and Precedent Retrieval) that can exploit the dependence between the two. We experiment extensively with several baseline models on the tasks, including lexical models, semantic models and ensemble based on GNNs. Further, to exploit the dependence between the two tasks, we develop an LLM-based re-ranking approach that gives the best performance.
format Preprint
id arxiv_https___arxiv_org_abs_2511_00268
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle IL-PCSR: Legal Corpus for Prior Case and Statute Retrieval
Paul, Shounak
Ghumare, Dhananjay
Goyal, Pawan
Ghosh, Saptarshi
Modi, Ashutosh
Computation and Language
Artificial Intelligence
Information Retrieval
Machine Learning
Identifying/retrieving relevant statutes and prior cases/precedents for a given legal situation are common tasks exercised by law practitioners. Researchers to date have addressed the two tasks independently, thus developing completely different datasets and models for each task; however, both retrieval tasks are inherently related, e.g., similar cases tend to cite similar statutes (due to similar factual situation). In this paper, we address this gap. We propose IL-PCR (Indian Legal corpus for Prior Case and Statute Retrieval), which is a unique corpus that provides a common testbed for developing models for both the tasks (Statute Retrieval and Precedent Retrieval) that can exploit the dependence between the two. We experiment extensively with several baseline models on the tasks, including lexical models, semantic models and ensemble based on GNNs. Further, to exploit the dependence between the two tasks, we develop an LLM-based re-ranking approach that gives the best performance.
title IL-PCSR: Legal Corpus for Prior Case and Statute Retrieval
topic Computation and Language
Artificial Intelligence
Information Retrieval
Machine Learning
url https://arxiv.org/abs/2511.00268