German Text Embedding Clustering Benchmark

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wehrli, Silvan, Arnrich, Bert, Irrgang, Christopher
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917559501062144
author Wehrli, Silvan
Arnrich, Bert
Irrgang, Christopher
author_facet Wehrli, Silvan
Arnrich, Bert
Irrgang, Christopher
contents This work introduces a benchmark assessing the performance of clustering German text embeddings in different domains. This benchmark is driven by the increasing use of clustering neural text embeddings in tasks that require the grouping of texts (such as topic modeling) and the need for German resources in existing benchmarks. We provide an initial analysis for a range of pre-trained mono- and multilingual models evaluated on the outcome of different clustering algorithms. Results include strong performing mono- and multilingual models. Reducing the dimensions of embeddings can further improve clustering. Additionally, we conduct experiments with continued pre-training for German BERT models to estimate the benefits of this additional training. Our experiments suggest that significant performance improvements are possible for short text. All code and datasets are publicly available.
format Preprint
id arxiv_https___arxiv_org_abs_2401_02709
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle German Text Embedding Clustering Benchmark
Wehrli, Silvan
Arnrich, Bert
Irrgang, Christopher
Computation and Language
Artificial Intelligence
This work introduces a benchmark assessing the performance of clustering German text embeddings in different domains. This benchmark is driven by the increasing use of clustering neural text embeddings in tasks that require the grouping of texts (such as topic modeling) and the need for German resources in existing benchmarks. We provide an initial analysis for a range of pre-trained mono- and multilingual models evaluated on the outcome of different clustering algorithms. Results include strong performing mono- and multilingual models. Reducing the dimensions of embeddings can further improve clustering. Additionally, we conduct experiments with continued pre-training for German BERT models to estimate the benefits of this additional training. Our experiments suggest that significant performance improvements are possible for short text. All code and datasets are publicly available.
title German Text Embedding Clustering Benchmark
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2401.02709