Neural Natural Language Processing for Long Texts: A Survey on Classification and Summarization

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Tsirmpas, Dimitrios, Gkionis, Ioannis, Papadopoulos, Georgios Th., Mademlis, Ioannis
Format: Preprint
Veröffentlicht: 2023
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866929276803088384
author Tsirmpas, Dimitrios
Gkionis, Ioannis
Papadopoulos, Georgios Th.
Mademlis, Ioannis
author_facet Tsirmpas, Dimitrios
Gkionis, Ioannis
Papadopoulos, Georgios Th.
Mademlis, Ioannis
contents The adoption of Deep Neural Networks (DNNs) has greatly benefited Natural Language Processing (NLP) during the past decade. However, the demands of long document analysis are quite different from those of shorter texts, while the ever increasing size of documents uploaded online renders automated understanding of lengthy texts a critical issue. Relevant applications include automated Web mining, legal document review, medical records analysis, financial reports analysis, contract management, environmental impact assessment, news aggregation, etc. Despite the relatively recent development of efficient algorithms for analyzing long documents, practical tools in this field are currently flourishing. This article serves as an entry point into this dynamic domain and aims to achieve two objectives. First of all, it provides an introductory overview of the relevant neural building blocks, serving as a concise tutorial for the field. Secondly, it offers a brief examination of the current state-of-the-art in two key long document analysis tasks: document classification and document summarization. Sentiment analysis for long texts is also covered, since it is typically treated as a particular case of document classification. Consequently, this article presents an introductory exploration of document-level analysis, addressing the primary challenges, concerns, and existing solutions. Finally, it offers a concise definition of "long text/document", presents an original overarching taxonomy of common deep neural methods for long document analysis and lists publicly available annotated datasets that can facilitate further research in this area.
format Preprint
id arxiv_https___arxiv_org_abs_2305_16259
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Neural Natural Language Processing for Long Texts: A Survey on Classification and Summarization
Tsirmpas, Dimitrios
Gkionis, Ioannis
Papadopoulos, Georgios Th.
Mademlis, Ioannis
Computation and Language
Artificial Intelligence
I.2.7
The adoption of Deep Neural Networks (DNNs) has greatly benefited Natural Language Processing (NLP) during the past decade. However, the demands of long document analysis are quite different from those of shorter texts, while the ever increasing size of documents uploaded online renders automated understanding of lengthy texts a critical issue. Relevant applications include automated Web mining, legal document review, medical records analysis, financial reports analysis, contract management, environmental impact assessment, news aggregation, etc. Despite the relatively recent development of efficient algorithms for analyzing long documents, practical tools in this field are currently flourishing. This article serves as an entry point into this dynamic domain and aims to achieve two objectives. First of all, it provides an introductory overview of the relevant neural building blocks, serving as a concise tutorial for the field. Secondly, it offers a brief examination of the current state-of-the-art in two key long document analysis tasks: document classification and document summarization. Sentiment analysis for long texts is also covered, since it is typically treated as a particular case of document classification. Consequently, this article presents an introductory exploration of document-level analysis, addressing the primary challenges, concerns, and existing solutions. Finally, it offers a concise definition of "long text/document", presents an original overarching taxonomy of common deep neural methods for long document analysis and lists publicly available annotated datasets that can facilitate further research in this area.
title Neural Natural Language Processing for Long Texts: A Survey on Classification and Summarization
topic Computation and Language
Artificial Intelligence
I.2.7
url https://arxiv.org/abs/2305.16259