Graph-Convolutional Networks: Named Entity Recognition and Large Language Model Embedding in Document Clustering

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Keraghel, Imed, Nadif, Mohamed
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913619006980096
author Keraghel, Imed
Nadif, Mohamed
author_facet Keraghel, Imed
Nadif, Mohamed
contents Recent advances in machine learning, particularly Large Language Models (LLMs) such as BERT and GPT, provide rich contextual embeddings that improve text representation. However, current document clustering approaches often ignore the deeper relationships between named entities (NEs) and the potential of LLM embeddings. This paper proposes a novel approach that integrates Named Entity Recognition (NER) and LLM embeddings within a graph-based framework for document clustering. The method builds a graph with nodes representing documents and edges weighted by named entity similarity, optimized using a graph-convolutional network (GCN). This ensures a more effective grouping of semantically related documents. Experimental results indicate that our approach outperforms conventional co-occurrence-based methods in clustering, notably for documents rich in named entities.
format Preprint
id arxiv_https___arxiv_org_abs_2412_14867
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Graph-Convolutional Networks: Named Entity Recognition and Large Language Model Embedding in Document Clustering
Keraghel, Imed
Nadif, Mohamed
Computation and Language
Recent advances in machine learning, particularly Large Language Models (LLMs) such as BERT and GPT, provide rich contextual embeddings that improve text representation. However, current document clustering approaches often ignore the deeper relationships between named entities (NEs) and the potential of LLM embeddings. This paper proposes a novel approach that integrates Named Entity Recognition (NER) and LLM embeddings within a graph-based framework for document clustering. The method builds a graph with nodes representing documents and edges weighted by named entity similarity, optimized using a graph-convolutional network (GCN). This ensures a more effective grouping of semantically related documents. Experimental results indicate that our approach outperforms conventional co-occurrence-based methods in clustering, notably for documents rich in named entities.
title Graph-Convolutional Networks: Named Entity Recognition and Large Language Model Embedding in Document Clustering
topic Computation and Language
url https://arxiv.org/abs/2412.14867