ERASMO: Leveraging Large Language Models for Enhanced Clustering Segmentation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Silva, Fillipe dos Santos, Kakimoto, Gabriel Kenzo, Reis, Julio Cesar dos, Reis, Marcelo S.
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866929698016067584
author Silva, Fillipe dos Santos
Kakimoto, Gabriel Kenzo
Reis, Julio Cesar dos
Reis, Marcelo S.
author_facet Silva, Fillipe dos Santos
Kakimoto, Gabriel Kenzo
Reis, Julio Cesar dos
Reis, Marcelo S.
contents Cluster analysis plays a crucial role in various domains and applications, such as customer segmentation in marketing. These contexts often involve multimodal data, including both tabular and textual datasets, making it challenging to represent hidden patterns for obtaining meaningful clusters. This study introduces ERASMO, a framework designed to fine-tune a pretrained language model on textually encoded tabular data and generate embeddings from the fine-tuned model. ERASMO employs a textual converter to transform tabular data into a textual format, enabling the language model to process and understand the data more effectively. Additionally, ERASMO produces contextually rich and structurally representative embeddings through techniques such as random feature sequence shuffling and number verbalization. Extensive experimental evaluations were conducted using multiple datasets and baseline approaches. Our results demonstrate that ERASMO fully leverages the specific context of each tabular dataset, leading to more precise and nuanced embeddings for accurate clustering. This approach enhances clustering performance by capturing complex relationship patterns within diverse tabular data.
format Preprint
id arxiv_https___arxiv_org_abs_2410_03738
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ERASMO: Leveraging Large Language Models for Enhanced Clustering Segmentation
Silva, Fillipe dos Santos
Kakimoto, Gabriel Kenzo
Reis, Julio Cesar dos
Reis, Marcelo S.
Computation and Language
Artificial Intelligence
68T50 (Natural language processing), 68T01 (General topics in artificial intelligence)
Cluster analysis plays a crucial role in various domains and applications, such as customer segmentation in marketing. These contexts often involve multimodal data, including both tabular and textual datasets, making it challenging to represent hidden patterns for obtaining meaningful clusters. This study introduces ERASMO, a framework designed to fine-tune a pretrained language model on textually encoded tabular data and generate embeddings from the fine-tuned model. ERASMO employs a textual converter to transform tabular data into a textual format, enabling the language model to process and understand the data more effectively. Additionally, ERASMO produces contextually rich and structurally representative embeddings through techniques such as random feature sequence shuffling and number verbalization. Extensive experimental evaluations were conducted using multiple datasets and baseline approaches. Our results demonstrate that ERASMO fully leverages the specific context of each tabular dataset, leading to more precise and nuanced embeddings for accurate clustering. This approach enhances clustering performance by capturing complex relationship patterns within diverse tabular data.
title ERASMO: Leveraging Large Language Models for Enhanced Clustering Segmentation
topic Computation and Language
Artificial Intelligence
68T50 (Natural language processing), 68T01 (General topics in artificial intelligence)
url https://arxiv.org/abs/2410.03738