LEGAL-UQA: A Low-Resource Urdu-English Dataset for Legal Question Answering

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Faisal, Faizan, Yousaf, Umair
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913550299037696
author Faisal, Faizan
Yousaf, Umair
author_facet Faisal, Faizan
Yousaf, Umair
contents We present LEGAL-UQA, the first Urdu legal question-answering dataset derived from Pakistan's constitution. This parallel English-Urdu dataset includes 619 question-answer pairs, each with corresponding legal article contexts, addressing the need for domain-specific NLP resources in low-resource languages. We describe the dataset creation process, including OCR extraction, manual refinement, and GPT-4-assisted translation and generation of QA pairs. Our experiments evaluate the latest generalist language and embedding models on LEGAL-UQA, with Claude-3.5-Sonnet achieving 99.19% human-evaluated accuracy. We fine-tune mt5-large-UQA-1.0, highlighting the challenges of adapting multilingual models to specialized domains. Additionally, we assess retrieval performance, finding OpenAI's text-embedding-3-large outperforms Mistral's mistral-embed. LEGAL-UQA bridges the gap between global NLP advancements and localized applications, particularly in constitutional law, and lays the foundation for improved legal information access in Pakistan.
format Preprint
id arxiv_https___arxiv_org_abs_2410_13013
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle LEGAL-UQA: A Low-Resource Urdu-English Dataset for Legal Question Answering
Faisal, Faizan
Yousaf, Umair
Computation and Language
Artificial Intelligence
Machine Learning
68T50
We present LEGAL-UQA, the first Urdu legal question-answering dataset derived from Pakistan's constitution. This parallel English-Urdu dataset includes 619 question-answer pairs, each with corresponding legal article contexts, addressing the need for domain-specific NLP resources in low-resource languages. We describe the dataset creation process, including OCR extraction, manual refinement, and GPT-4-assisted translation and generation of QA pairs. Our experiments evaluate the latest generalist language and embedding models on LEGAL-UQA, with Claude-3.5-Sonnet achieving 99.19% human-evaluated accuracy. We fine-tune mt5-large-UQA-1.0, highlighting the challenges of adapting multilingual models to specialized domains. Additionally, we assess retrieval performance, finding OpenAI's text-embedding-3-large outperforms Mistral's mistral-embed. LEGAL-UQA bridges the gap between global NLP advancements and localized applications, particularly in constitutional law, and lays the foundation for improved legal information access in Pakistan.
title LEGAL-UQA: A Low-Resource Urdu-English Dataset for Legal Question Answering
topic Computation and Language
Artificial Intelligence
Machine Learning
68T50
url https://arxiv.org/abs/2410.13013