A Reasoning-Focused Legal Retrieval Benchmark

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zheng, Lucia, Guha, Neel, Arifov, Javokhir, Zhang, Sarah, Skreta, Michal, Manning, Christopher D., Henderson, Peter, Ho, Daniel E.
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912363395940352
author Zheng, Lucia
Guha, Neel
Arifov, Javokhir
Zhang, Sarah
Skreta, Michal
Manning, Christopher D.
Henderson, Peter
Ho, Daniel E.
author_facet Zheng, Lucia
Guha, Neel
Arifov, Javokhir
Zhang, Sarah
Skreta, Michal
Manning, Christopher D.
Henderson, Peter
Ho, Daniel E.
contents As the legal community increasingly examines the use of large language models (LLMs) for various legal applications, legal AI developers have turned to retrieval-augmented LLMs ("RAG" systems) to improve system performance and robustness. An obstacle to the development of specialized RAG systems is the lack of realistic legal RAG benchmarks which capture the complexity of both legal retrieval and downstream legal question-answering. To address this, we introduce two novel legal RAG benchmarks: Bar Exam QA and Housing Statute QA. Our tasks correspond to real-world legal research tasks, and were produced through annotation processes which resemble legal research. We describe the construction of these benchmarks and the performance of existing retriever pipelines. Our results suggest that legal RAG remains a challenging application, thus motivating future research.
format Preprint
id arxiv_https___arxiv_org_abs_2505_03970
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Reasoning-Focused Legal Retrieval Benchmark
Zheng, Lucia
Guha, Neel
Arifov, Javokhir
Zhang, Sarah
Skreta, Michal
Manning, Christopher D.
Henderson, Peter
Ho, Daniel E.
Computation and Language
As the legal community increasingly examines the use of large language models (LLMs) for various legal applications, legal AI developers have turned to retrieval-augmented LLMs ("RAG" systems) to improve system performance and robustness. An obstacle to the development of specialized RAG systems is the lack of realistic legal RAG benchmarks which capture the complexity of both legal retrieval and downstream legal question-answering. To address this, we introduce two novel legal RAG benchmarks: Bar Exam QA and Housing Statute QA. Our tasks correspond to real-world legal research tasks, and were produced through annotation processes which resemble legal research. We describe the construction of these benchmarks and the performance of existing retriever pipelines. Our results suggest that legal RAG remains a challenging application, thus motivating future research.
title A Reasoning-Focused Legal Retrieval Benchmark
topic Computation and Language
url https://arxiv.org/abs/2505.03970