Saved in:
Bibliographic Details
Main Authors: Keuren, Paul, Ponsen, Marc, Bagheri, Robert Ayoub
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2604.21555
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908989466345472
author Keuren, Paul
Ponsen, Marc
Bagheri, Robert Ayoub
author_facet Keuren, Paul
Ponsen, Marc
Bagheri, Robert Ayoub
contents Sentence embedding techniques aim to encode key concepts of a sentence's meaning in a vector space. However, the majority of evaluation approaches for sentence embedding quality rely on the use of additional classifiers or downstream tasks. These additional components make it unclear whether good results stem from the embedding itself or from the classifier's behaviour. In this paper, we propose a novel method for evaluating the effectiveness of sentence embedding methods in capturing sentence-level concepts. Our approach is classifier-independent, allowing for an objective assessment of the model's performance. The approach adopted in this study involves the systematic introduction of syntactic noise and semantic negations into sentences, with the subsequent quantification of their relative effects on the resulting embeddings. The visualisation of these effects is facilitated by Concept Separation Curves, which show the model's capacity to differentiate between conceptual and surface-level variations. By leveraging data from multiple domains, employing both Dutch and English languages, and examining sentence lengths, this study offers a compelling demonstration that Concept Separation Curves provide an interpretable, reproducible, and cross-model approach for evaluating the conceptual stability of sentence embeddings.
format Preprint
id arxiv_https___arxiv_org_abs_2604_21555
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Finding Meaning in Embeddings: Concept Separation Curves
Keuren, Paul
Ponsen, Marc
Bagheri, Robert Ayoub
Computation and Language
Sentence embedding techniques aim to encode key concepts of a sentence's meaning in a vector space. However, the majority of evaluation approaches for sentence embedding quality rely on the use of additional classifiers or downstream tasks. These additional components make it unclear whether good results stem from the embedding itself or from the classifier's behaviour. In this paper, we propose a novel method for evaluating the effectiveness of sentence embedding methods in capturing sentence-level concepts. Our approach is classifier-independent, allowing for an objective assessment of the model's performance. The approach adopted in this study involves the systematic introduction of syntactic noise and semantic negations into sentences, with the subsequent quantification of their relative effects on the resulting embeddings. The visualisation of these effects is facilitated by Concept Separation Curves, which show the model's capacity to differentiate between conceptual and surface-level variations. By leveraging data from multiple domains, employing both Dutch and English languages, and examining sentence lengths, this study offers a compelling demonstration that Concept Separation Curves provide an interpretable, reproducible, and cross-model approach for evaluating the conceptual stability of sentence embeddings.
title Finding Meaning in Embeddings: Concept Separation Curves
topic Computation and Language
url https://arxiv.org/abs/2604.21555