Fine-tuning the SwissBERT Encoder Model for Embedding Sentences and Documents

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Grosjean, Juri, Vamvas, Jannis
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914793833627648
author Grosjean, Juri
Vamvas, Jannis
author_facet Grosjean, Juri
Vamvas, Jannis
contents Encoder models trained for the embedding of sentences or short documents have proven useful for tasks such as semantic search and topic modeling. In this paper, we present a version of the SwissBERT encoder model that we specifically fine-tuned for this purpose. SwissBERT contains language adapters for the four national languages of Switzerland -- German, French, Italian, and Romansh -- and has been pre-trained on a large number of news articles in those languages. Using contrastive learning based on a subset of these articles, we trained a fine-tuned version, which we call SentenceSwissBERT. Multilingual experiments on document retrieval and text classification in a Switzerland-specific setting show that SentenceSwissBERT surpasses the accuracy of the original SwissBERT model and of a comparable baseline. The model is openly available for research use.
format Preprint
id arxiv_https___arxiv_org_abs_2405_07513
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Fine-tuning the SwissBERT Encoder Model for Embedding Sentences and Documents
Grosjean, Juri
Vamvas, Jannis
Computation and Language
Encoder models trained for the embedding of sentences or short documents have proven useful for tasks such as semantic search and topic modeling. In this paper, we present a version of the SwissBERT encoder model that we specifically fine-tuned for this purpose. SwissBERT contains language adapters for the four national languages of Switzerland -- German, French, Italian, and Romansh -- and has been pre-trained on a large number of news articles in those languages. Using contrastive learning based on a subset of these articles, we trained a fine-tuned version, which we call SentenceSwissBERT. Multilingual experiments on document retrieval and text classification in a Switzerland-specific setting show that SentenceSwissBERT surpasses the accuracy of the original SwissBERT model and of a comparable baseline. The model is openly available for research use.
title Fine-tuning the SwissBERT Encoder Model for Embedding Sentences and Documents
topic Computation and Language
url https://arxiv.org/abs/2405.07513