Space Decomposition for Sentence Embedding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ponwitayarat, Wuttikorn, Limkonchotiwat, Peerat, Chuangsuwanich, Ekapol, Nutanong, Sarana
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914824829534208
author Ponwitayarat, Wuttikorn
Limkonchotiwat, Peerat
Chuangsuwanich, Ekapol
Nutanong, Sarana
author_facet Ponwitayarat, Wuttikorn
Limkonchotiwat, Peerat
Chuangsuwanich, Ekapol
Nutanong, Sarana
contents Determining sentence pair similarity is crucial for various NLP tasks. A common technique to address this is typically evaluated on a continuous semantic textual similarity scale from 0 to 5. However, based on a linguistic observation in STS annotation guidelines, we found that the score in the range [4,5] indicates an upper-range sample, while the rest are lower-range samples. This necessitates a new approach to treating the upper-range and lower-range classes separately. In this paper, we introduce a novel embedding space decomposition method called MixSP utilizing a Mixture of Specialized Projectors, designed to distinguish and rank upper-range and lower-range samples accurately. The experimental results demonstrate that MixSP decreased the overlap representation between upper-range and lower-range classes significantly while outperforming competitors on STS and zero-shot benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2406_03125
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Space Decomposition for Sentence Embedding
Ponwitayarat, Wuttikorn
Limkonchotiwat, Peerat
Chuangsuwanich, Ekapol
Nutanong, Sarana
Computation and Language
Determining sentence pair similarity is crucial for various NLP tasks. A common technique to address this is typically evaluated on a continuous semantic textual similarity scale from 0 to 5. However, based on a linguistic observation in STS annotation guidelines, we found that the score in the range [4,5] indicates an upper-range sample, while the rest are lower-range samples. This necessitates a new approach to treating the upper-range and lower-range classes separately. In this paper, we introduce a novel embedding space decomposition method called MixSP utilizing a Mixture of Specialized Projectors, designed to distinguish and rank upper-range and lower-range samples accurately. The experimental results demonstrate that MixSP decreased the overlap representation between upper-range and lower-range classes significantly while outperforming competitors on STS and zero-shot benchmarks.
title Space Decomposition for Sentence Embedding
topic Computation and Language
url https://arxiv.org/abs/2406.03125