The Cambridge Law Corpus: A Dataset for Legal AI Research

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Östling, Andreas, Sargeant, Holli, Xie, Huiyuan, Bull, Ludwig, Terenin, Alexander, Jonsson, Leif, Magnusson, Måns, Steffek, Felix
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916079225274368
author Östling, Andreas
Sargeant, Holli
Xie, Huiyuan
Bull, Ludwig
Terenin, Alexander
Jonsson, Leif
Magnusson, Måns
Steffek, Felix
author_facet Östling, Andreas
Sargeant, Holli
Xie, Huiyuan
Bull, Ludwig
Terenin, Alexander
Jonsson, Leif
Magnusson, Måns
Steffek, Felix
contents We introduce the Cambridge Law Corpus (CLC), a dataset for legal AI research. It consists of over 250 000 court cases from the UK. Most cases are from the 21st century, but the corpus includes cases as old as the 16th century. This paper presents the first release of the corpus, containing the raw text and meta-data. Together with the corpus, we provide annotations on case outcomes for 638 cases, done by legal experts. Using our annotated data, we have trained and evaluated case outcome extraction with GPT-3, GPT-4 and RoBERTa models to provide benchmarks. We include an extensive legal and ethical discussion to address the potentially sensitive nature of this material. As a consequence, the corpus will only be released for research purposes under certain restrictions.
format Preprint
id arxiv_https___arxiv_org_abs_2309_12269
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle The Cambridge Law Corpus: A Dataset for Legal AI Research
Östling, Andreas
Sargeant, Holli
Xie, Huiyuan
Bull, Ludwig
Terenin, Alexander
Jonsson, Leif
Magnusson, Måns
Steffek, Felix
Computation and Language
Computers and Society
Applications
We introduce the Cambridge Law Corpus (CLC), a dataset for legal AI research. It consists of over 250 000 court cases from the UK. Most cases are from the 21st century, but the corpus includes cases as old as the 16th century. This paper presents the first release of the corpus, containing the raw text and meta-data. Together with the corpus, we provide annotations on case outcomes for 638 cases, done by legal experts. Using our annotated data, we have trained and evaluated case outcome extraction with GPT-3, GPT-4 and RoBERTa models to provide benchmarks. We include an extensive legal and ethical discussion to address the potentially sensitive nature of this material. As a consequence, the corpus will only be released for research purposes under certain restrictions.
title The Cambridge Law Corpus: A Dataset for Legal AI Research
topic Computation and Language
Computers and Society
Applications
url https://arxiv.org/abs/2309.12269