No Argument Left Behind: Overlapping Chunks for Faster Processing of Arbitrarily Long Legal Texts
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909429655404544 |
|---|---|
| author | Fama, Israel Bueno, Bárbara Alcoforado, Alexandre Ferraz, Thomas Palmeira Moya, Arnold Costa, Anna Helena Reali |
| author_facet | Fama, Israel Bueno, Bárbara Alcoforado, Alexandre Ferraz, Thomas Palmeira Moya, Arnold Costa, Anna Helena Reali |
| contents | In a context where the Brazilian judiciary system, the largest in the world, faces a crisis due to the slow processing of millions of cases, it becomes imperative to develop efficient methods for analyzing legal texts. We introduce uBERT, a hybrid model that combines Transformer and Recurrent Neural Network architectures to effectively handle long legal texts. Our approach processes the full text regardless of its length while maintaining reasonable computational overhead. Our experiments demonstrate that uBERT achieves superior performance compared to BERT+LSTM when overlapping input is used and is significantly faster than ULMFiT for processing long legal documents. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2410_19184 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | No Argument Left Behind: Overlapping Chunks for Faster Processing of Arbitrarily Long Legal Texts Fama, Israel Bueno, Bárbara Alcoforado, Alexandre Ferraz, Thomas Palmeira Moya, Arnold Costa, Anna Helena Reali Computation and Language Artificial Intelligence Computers and Society Machine Learning In a context where the Brazilian judiciary system, the largest in the world, faces a crisis due to the slow processing of millions of cases, it becomes imperative to develop efficient methods for analyzing legal texts. We introduce uBERT, a hybrid model that combines Transformer and Recurrent Neural Network architectures to effectively handle long legal texts. Our approach processes the full text regardless of its length while maintaining reasonable computational overhead. Our experiments demonstrate that uBERT achieves superior performance compared to BERT+LSTM when overlapping input is used and is significantly faster than ULMFiT for processing long legal documents. |
| title | No Argument Left Behind: Overlapping Chunks for Faster Processing of Arbitrarily Long Legal Texts |
| topic | Computation and Language Artificial Intelligence Computers and Society Machine Learning |
| url | https://arxiv.org/abs/2410.19184 |