An Improved Method for Class-specific Keyword Extraction: A Case Study in the German Business Registry

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Meisenbacher, Stephen, Schopf, Tim, Yan, Weixin, Holl, Patrick, Matthes, Florian
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910535062126592
author Meisenbacher, Stephen
Schopf, Tim
Yan, Weixin
Holl, Patrick
Matthes, Florian
author_facet Meisenbacher, Stephen
Schopf, Tim
Yan, Weixin
Holl, Patrick
Matthes, Florian
contents The task of $\textit{keyword extraction}$ is often an important initial step in unsupervised information extraction, forming the basis for tasks such as topic modeling or document classification. While recent methods have proven to be quite effective in the extraction of keywords, the identification of $\textit{class-specific}$ keywords, or only those pertaining to a predefined class, remains challenging. In this work, we propose an improved method for class-specific keyword extraction, which builds upon the popular $\textbf{KeyBERT}$ library to identify only keywords related to a class described by $\textit{seed keywords}$. We test this method using a dataset of German business registry entries, where the goal is to classify each business according to an economic sector. Our results reveal that our method greatly improves upon previous approaches, setting a new standard for $\textit{class-specific}$ keyword extraction.
format Preprint
id arxiv_https___arxiv_org_abs_2407_14085
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle An Improved Method for Class-specific Keyword Extraction: A Case Study in the German Business Registry
Meisenbacher, Stephen
Schopf, Tim
Yan, Weixin
Holl, Patrick
Matthes, Florian
Computation and Language
The task of $\textit{keyword extraction}$ is often an important initial step in unsupervised information extraction, forming the basis for tasks such as topic modeling or document classification. While recent methods have proven to be quite effective in the extraction of keywords, the identification of $\textit{class-specific}$ keywords, or only those pertaining to a predefined class, remains challenging. In this work, we propose an improved method for class-specific keyword extraction, which builds upon the popular $\textbf{KeyBERT}$ library to identify only keywords related to a class described by $\textit{seed keywords}$. We test this method using a dataset of German business registry entries, where the goal is to classify each business according to an economic sector. Our results reveal that our method greatly improves upon previous approaches, setting a new standard for $\textit{class-specific}$ keyword extraction.
title An Improved Method for Class-specific Keyword Extraction: A Case Study in the German Business Registry
topic Computation and Language
url https://arxiv.org/abs/2407.14085