Structured Definitions and Segmentations for Legal Reasoning in LLMs: A Study on Indian Legal Data

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Khatri, Mann, Yusuf, Mirza, Shah, Rajiv Ratn, Kumaraguru, Ponnurangam
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912729391955968
author Khatri, Mann
Yusuf, Mirza
Shah, Rajiv Ratn
Kumaraguru, Ponnurangam
author_facet Khatri, Mann
Yusuf, Mirza
Shah, Rajiv Ratn
Kumaraguru, Ponnurangam
contents Large Language Models (LLMs), trained on extensive datasets from the web, exhibit remarkable general reasoning skills. Despite this, they often struggle in specialized areas like law, mainly because they lack domain-specific pretraining. The legal field presents unique challenges, as legal documents are generally long and intricate, making it hard for models to process the full text efficiently. Previous studies have examined in-context approaches to address the knowledge gap, boosting model performance in new domains without full domain alignment. In our paper, we analyze model behavior on legal tasks by conducting experiments in three areas: (i) reorganizing documents based on rhetorical roles to assess how structured information affects long context processing and model decisions, (ii) defining rhetorical roles to familiarize the model with legal terminology, and (iii) emulating the step-by-step reasoning of courts regarding rhetorical roles to enhance model reasoning. These experiments are conducted in a zero-shot setting across three Indian legal judgment prediction datasets. Our results reveal that organizing data or explaining key legal terms significantly boosts model performance, with a minimum increase of ~1.5% and a maximum improvement of 4.36% in F1 score compared to the baseline.
format Preprint
id arxiv_https___arxiv_org_abs_2511_20669
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Structured Definitions and Segmentations for Legal Reasoning in LLMs: A Study on Indian Legal Data
Khatri, Mann
Yusuf, Mirza
Shah, Rajiv Ratn
Kumaraguru, Ponnurangam
Computation and Language
Artificial Intelligence
Large Language Models (LLMs), trained on extensive datasets from the web, exhibit remarkable general reasoning skills. Despite this, they often struggle in specialized areas like law, mainly because they lack domain-specific pretraining. The legal field presents unique challenges, as legal documents are generally long and intricate, making it hard for models to process the full text efficiently. Previous studies have examined in-context approaches to address the knowledge gap, boosting model performance in new domains without full domain alignment. In our paper, we analyze model behavior on legal tasks by conducting experiments in three areas: (i) reorganizing documents based on rhetorical roles to assess how structured information affects long context processing and model decisions, (ii) defining rhetorical roles to familiarize the model with legal terminology, and (iii) emulating the step-by-step reasoning of courts regarding rhetorical roles to enhance model reasoning. These experiments are conducted in a zero-shot setting across three Indian legal judgment prediction datasets. Our results reveal that organizing data or explaining key legal terms significantly boosts model performance, with a minimum increase of ~1.5% and a maximum improvement of 4.36% in F1 score compared to the baseline.
title Structured Definitions and Segmentations for Legal Reasoning in LLMs: A Study on Indian Legal Data
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2511.20669