Self-Supervised Borrowing Detection on Multilingual Wordlists

Fuente: arXiv
Saved in:
Bibliographic Details
Main Author: Wientzek, Tim
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912740904271872
author Wientzek, Tim
author_facet Wientzek, Tim
contents This paper presents a fully self-supervised approach to borrowing detection in multilingual wordlists. The method combines two sources of information: PMI similarities based on a global correspondence model and a lightweight contrastive component trained on phonetic feature vectors. It further includes an automatic procedure for selecting decision thresholds without requiring labeled data. Experiments on benchmark datasets show that PMI alone already improves over existing string similarity measures such as NED and SCA, and that the combined similarity performs on par with or better than supervised baselines. An ablation study highlights the importance of character encoding, temperature settings and augmentation strategies. The approach scales to datasets of different sizes, works without manual supervision and is provided with a command-line tool that allows researchers to conduct their own studies.
format Preprint
id arxiv_https___arxiv_org_abs_2512_01713
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Self-Supervised Borrowing Detection on Multilingual Wordlists
Wientzek, Tim
Computation and Language
J.5; I.2.7; I.5.0
This paper presents a fully self-supervised approach to borrowing detection in multilingual wordlists. The method combines two sources of information: PMI similarities based on a global correspondence model and a lightweight contrastive component trained on phonetic feature vectors. It further includes an automatic procedure for selecting decision thresholds without requiring labeled data. Experiments on benchmark datasets show that PMI alone already improves over existing string similarity measures such as NED and SCA, and that the combined similarity performs on par with or better than supervised baselines. An ablation study highlights the importance of character encoding, temperature settings and augmentation strategies. The approach scales to datasets of different sizes, works without manual supervision and is provided with a command-line tool that allows researchers to conduct their own studies.
title Self-Supervised Borrowing Detection on Multilingual Wordlists
topic Computation and Language
J.5; I.2.7; I.5.0
url https://arxiv.org/abs/2512.01713