Discovering Significant Topics from Legal Decisions with Selective Inference

Fuente: arXiv
Saved in:
Bibliographic Details
Main Author: Soh, Jerrold
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916079476932608
author Soh, Jerrold
author_facet Soh, Jerrold
contents We propose and evaluate an automated pipeline for discovering significant topics from legal decision texts by passing features synthesized with topic models through penalised regressions and post-selection significance tests. The method identifies case topics significantly correlated with outcomes, topic-word distributions which can be manually-interpreted to gain insights about significant topics, and case-topic weights which can be used to identify representative cases for each topic. We demonstrate the method on a new dataset of domain name disputes and a canonical dataset of European Court of Human Rights violation cases. Topic models based on latent semantic analysis as well as language model embeddings are evaluated. We show that topics derived by the pipeline are consistent with legal doctrines in both areas and can be useful in other related legal analysis tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2401_01068
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Discovering Significant Topics from Legal Decisions with Selective Inference
Soh, Jerrold
Computation and Language
Artificial Intelligence
We propose and evaluate an automated pipeline for discovering significant topics from legal decision texts by passing features synthesized with topic models through penalised regressions and post-selection significance tests. The method identifies case topics significantly correlated with outcomes, topic-word distributions which can be manually-interpreted to gain insights about significant topics, and case-topic weights which can be used to identify representative cases for each topic. We demonstrate the method on a new dataset of domain name disputes and a canonical dataset of European Court of Human Rights violation cases. Topic models based on latent semantic analysis as well as language model embeddings are evaluated. We show that topics derived by the pipeline are consistent with legal doctrines in both areas and can be useful in other related legal analysis tasks.
title Discovering Significant Topics from Legal Decisions with Selective Inference
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2401.01068