One Law, Many Languages: Benchmarking Multilingual Legal Reasoning for Judicial Support

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Stern, Ronja, Rasiah, Vishvaksenan, Matoshi, Veton, Bose, Srinanda Brügger, Stürmer, Matthias, Chalkidis, Ilias, Ho, Daniel E., Niklaus, Joel
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913474768011264
author Stern, Ronja
Rasiah, Vishvaksenan
Matoshi, Veton
Bose, Srinanda Brügger
Stürmer, Matthias
Chalkidis, Ilias
Ho, Daniel E.
Niklaus, Joel
author_facet Stern, Ronja
Rasiah, Vishvaksenan
Matoshi, Veton
Bose, Srinanda Brügger
Stürmer, Matthias
Chalkidis, Ilias
Ho, Daniel E.
Niklaus, Joel
contents Recent strides in Large Language Models (LLMs) have saturated many Natural Language Processing (NLP) benchmarks, emphasizing the need for more challenging ones to properly assess LLM capabilities. However, domain-specific and multilingual benchmarks are rare because they require in-depth expertise to develop. Still, most public models are trained predominantly on English corpora, while other languages remain understudied, particularly for practical domain-specific NLP tasks. In this work, we introduce a novel NLP benchmark for the legal domain that challenges LLMs in five key dimensions: processing \emph{long documents} (up to 50K tokens), using \emph{domain-specific knowledge} (embodied in legal texts), \emph{multilingual} understanding (covering five languages), \emph{multitasking} (comprising legal document-to-document Information Retrieval, Court View Generation, Leading Decision Summarization, Citation Extraction, and eight challenging Text Classification tasks) and \emph{reasoning} (comprising especially Court View Generation, but also the Text Classification tasks). Our benchmark contains diverse datasets from the Swiss legal system, allowing for a comprehensive study of the underlying non-English, inherently multilingual legal system. Despite the large size of our datasets (some with hundreds of thousands of examples), existing publicly available multilingual models struggle with most tasks, even after extensive in-domain pre-training and fine-tuning. We publish all resources (benchmark suite, pre-trained models, code) under permissive open CC BY-SA licenses.
format Preprint
id arxiv_https___arxiv_org_abs_2306_09237
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle One Law, Many Languages: Benchmarking Multilingual Legal Reasoning for Judicial Support
Stern, Ronja
Rasiah, Vishvaksenan
Matoshi, Veton
Bose, Srinanda Brügger
Stürmer, Matthias
Chalkidis, Ilias
Ho, Daniel E.
Niklaus, Joel
Computation and Language
Artificial Intelligence
Machine Learning
68T50
I.2
Recent strides in Large Language Models (LLMs) have saturated many Natural Language Processing (NLP) benchmarks, emphasizing the need for more challenging ones to properly assess LLM capabilities. However, domain-specific and multilingual benchmarks are rare because they require in-depth expertise to develop. Still, most public models are trained predominantly on English corpora, while other languages remain understudied, particularly for practical domain-specific NLP tasks. In this work, we introduce a novel NLP benchmark for the legal domain that challenges LLMs in five key dimensions: processing \emph{long documents} (up to 50K tokens), using \emph{domain-specific knowledge} (embodied in legal texts), \emph{multilingual} understanding (covering five languages), \emph{multitasking} (comprising legal document-to-document Information Retrieval, Court View Generation, Leading Decision Summarization, Citation Extraction, and eight challenging Text Classification tasks) and \emph{reasoning} (comprising especially Court View Generation, but also the Text Classification tasks). Our benchmark contains diverse datasets from the Swiss legal system, allowing for a comprehensive study of the underlying non-English, inherently multilingual legal system. Despite the large size of our datasets (some with hundreds of thousands of examples), existing publicly available multilingual models struggle with most tasks, even after extensive in-domain pre-training and fine-tuning. We publish all resources (benchmark suite, pre-trained models, code) under permissive open CC BY-SA licenses.
title One Law, Many Languages: Benchmarking Multilingual Legal Reasoning for Judicial Support
topic Computation and Language
Artificial Intelligence
Machine Learning
68T50
I.2
url https://arxiv.org/abs/2306.09237