JUÁ -- A Benchmark for Information Retrieval in Brazilian Legal Text Collections

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Pereira, Jayr, Fernandes, Leandro, de Brito, Erick, Lotufo, Roberto, Bonifacio, Luiz
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918434342699008
author Pereira, Jayr
Fernandes, Leandro
de Brito, Erick
Lotufo, Roberto
Bonifacio, Luiz
author_facet Pereira, Jayr
Fernandes, Leandro
de Brito, Erick
Lotufo, Roberto
Bonifacio, Luiz
contents Legal information retrieval in Portuguese remains difficult to evaluate systematically because available datasets differ widely in document type, query style, and relevance definition. We present JUÁ, a public benchmark for Brazilian legal retrieval designed to support more reproducible and comparable evaluation across heterogeneous legal collections. More broadly, JUÁ is intended not only as a benchmark, but as a continuous evaluation infrastructure for Brazilian legal IR, combining shared protocols, common ranking metrics, fixed splits when applicable, and a public leaderboard. The benchmark covers jurisprudence retrieval as well as broader legislative, regulatory, and question-driven legal search. We evaluate lexical, dense, and BM25-based reranking pipelines, including a domain-adapted Qwen embedding model fine-tuned on JUÁ-aligned supervision. Results show that the benchmark is sufficiently heterogeneous to distinguish retrieval paradigms and reveal substantial cross-dataset trade-offs. Domain adaptation yields its clearest gains on the supervision-aligned JUÁ-Juris subset, while BM25 remains highly competitive on other collections, especially in settings with strong lexical and institutional phrasing cues. Overall, JUÁ provides a practical evaluation framework for studying legal retrieval across multiple Brazilian legal domains under a common benchmark design.
format Preprint
id arxiv_https___arxiv_org_abs_2604_06098
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle JUÁ -- A Benchmark for Information Retrieval in Brazilian Legal Text Collections
Pereira, Jayr
Fernandes, Leandro
de Brito, Erick
Lotufo, Roberto
Bonifacio, Luiz
Information Retrieval
Computation and Language
Legal information retrieval in Portuguese remains difficult to evaluate systematically because available datasets differ widely in document type, query style, and relevance definition. We present JUÁ, a public benchmark for Brazilian legal retrieval designed to support more reproducible and comparable evaluation across heterogeneous legal collections. More broadly, JUÁ is intended not only as a benchmark, but as a continuous evaluation infrastructure for Brazilian legal IR, combining shared protocols, common ranking metrics, fixed splits when applicable, and a public leaderboard. The benchmark covers jurisprudence retrieval as well as broader legislative, regulatory, and question-driven legal search. We evaluate lexical, dense, and BM25-based reranking pipelines, including a domain-adapted Qwen embedding model fine-tuned on JUÁ-aligned supervision. Results show that the benchmark is sufficiently heterogeneous to distinguish retrieval paradigms and reveal substantial cross-dataset trade-offs. Domain adaptation yields its clearest gains on the supervision-aligned JUÁ-Juris subset, while BM25 remains highly competitive on other collections, especially in settings with strong lexical and institutional phrasing cues. Overall, JUÁ provides a practical evaluation framework for studying legal retrieval across multiple Brazilian legal domains under a common benchmark design.
title JUÁ -- A Benchmark for Information Retrieval in Brazilian Legal Text Collections
topic Information Retrieval
Computation and Language
url https://arxiv.org/abs/2604.06098