LongKey: Keyphrase Extraction for Long Documents

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Alves, Jeovane Honorio, State, Radu, Freitas, Cinthia Obladen de Almendra, Barddal, Jean Paul
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929681471635456
author Alves, Jeovane Honorio
State, Radu
Freitas, Cinthia Obladen de Almendra
Barddal, Jean Paul
author_facet Alves, Jeovane Honorio
State, Radu
Freitas, Cinthia Obladen de Almendra
Barddal, Jean Paul
contents In an era of information overload, manually annotating the vast and growing corpus of documents and scholarly papers is increasingly impractical. Automated keyphrase extraction addresses this challenge by identifying representative terms within texts. However, most existing methods focus on short documents (up to 512 tokens), leaving a gap in processing long-context documents. In this paper, we introduce LongKey, a novel framework for extracting keyphrases from lengthy documents, which uses an encoder-based language model to capture extended text intricacies. LongKey uses a max-pooling embedder to enhance keyphrase candidate representation. Validated on the comprehensive LDKP datasets and six diverse, unseen datasets, LongKey consistently outperforms existing unsupervised and language model-based keyphrase extraction methods. Our findings demonstrate LongKey's versatility and superior performance, marking an advancement in keyphrase extraction for varied text lengths and domains.
format Preprint
id arxiv_https___arxiv_org_abs_2411_17863
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle LongKey: Keyphrase Extraction for Long Documents
Alves, Jeovane Honorio
State, Radu
Freitas, Cinthia Obladen de Almendra
Barddal, Jean Paul
Computation and Language
Artificial Intelligence
Information Retrieval
Machine Learning
In an era of information overload, manually annotating the vast and growing corpus of documents and scholarly papers is increasingly impractical. Automated keyphrase extraction addresses this challenge by identifying representative terms within texts. However, most existing methods focus on short documents (up to 512 tokens), leaving a gap in processing long-context documents. In this paper, we introduce LongKey, a novel framework for extracting keyphrases from lengthy documents, which uses an encoder-based language model to capture extended text intricacies. LongKey uses a max-pooling embedder to enhance keyphrase candidate representation. Validated on the comprehensive LDKP datasets and six diverse, unseen datasets, LongKey consistently outperforms existing unsupervised and language model-based keyphrase extraction methods. Our findings demonstrate LongKey's versatility and superior performance, marking an advancement in keyphrase extraction for varied text lengths and domains.
title LongKey: Keyphrase Extraction for Long Documents
topic Computation and Language
Artificial Intelligence
Information Retrieval
Machine Learning
url https://arxiv.org/abs/2411.17863