No Argument Left Behind: Overlapping Chunks for Faster Processing of Arbitrarily Long Legal Texts

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fama, Israel, Bueno, Bárbara, Alcoforado, Alexandre, Ferraz, Thomas Palmeira, Moya, Arnold, Costa, Anna Helena Reali
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909429655404544
author Fama, Israel
Bueno, Bárbara
Alcoforado, Alexandre
Ferraz, Thomas Palmeira
Moya, Arnold
Costa, Anna Helena Reali
author_facet Fama, Israel
Bueno, Bárbara
Alcoforado, Alexandre
Ferraz, Thomas Palmeira
Moya, Arnold
Costa, Anna Helena Reali
contents In a context where the Brazilian judiciary system, the largest in the world, faces a crisis due to the slow processing of millions of cases, it becomes imperative to develop efficient methods for analyzing legal texts. We introduce uBERT, a hybrid model that combines Transformer and Recurrent Neural Network architectures to effectively handle long legal texts. Our approach processes the full text regardless of its length while maintaining reasonable computational overhead. Our experiments demonstrate that uBERT achieves superior performance compared to BERT+LSTM when overlapping input is used and is significantly faster than ULMFiT for processing long legal documents.
format Preprint
id arxiv_https___arxiv_org_abs_2410_19184
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle No Argument Left Behind: Overlapping Chunks for Faster Processing of Arbitrarily Long Legal Texts
Fama, Israel
Bueno, Bárbara
Alcoforado, Alexandre
Ferraz, Thomas Palmeira
Moya, Arnold
Costa, Anna Helena Reali
Computation and Language
Artificial Intelligence
Computers and Society
Machine Learning
In a context where the Brazilian judiciary system, the largest in the world, faces a crisis due to the slow processing of millions of cases, it becomes imperative to develop efficient methods for analyzing legal texts. We introduce uBERT, a hybrid model that combines Transformer and Recurrent Neural Network architectures to effectively handle long legal texts. Our approach processes the full text regardless of its length while maintaining reasonable computational overhead. Our experiments demonstrate that uBERT achieves superior performance compared to BERT+LSTM when overlapping input is used and is significantly faster than ULMFiT for processing long legal documents.
title No Argument Left Behind: Overlapping Chunks for Faster Processing of Arbitrarily Long Legal Texts
topic Computation and Language
Artificial Intelligence
Computers and Society
Machine Learning
url https://arxiv.org/abs/2410.19184