Subspace Representations for Soft Set Operations and Sentence Similarities

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ishibashi, Yoichi, Yokoi, Sho, Sudoh, Katsuhito, Nakamura, Satoshi
Format: Preprint
Published: 2022
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916198636060672
author Ishibashi, Yoichi
Yokoi, Sho
Sudoh, Katsuhito
Nakamura, Satoshi
author_facet Ishibashi, Yoichi
Yokoi, Sho
Sudoh, Katsuhito
Nakamura, Satoshi
contents In the field of natural language processing (NLP), continuous vector representations are crucial for capturing the semantic meanings of individual words. Yet, when it comes to the representations of sets of words, the conventional vector-based approaches often struggle with expressiveness and lack the essential set operations such as union, intersection, and complement. Inspired by quantum logic, we realize the representation of word sets and corresponding set operations within pre-trained word embedding spaces. By grounding our approach in the linear subspaces, we enable efficient computation of various set operations and facilitate the soft computation of membership functions within continuous spaces. Moreover, we allow for the computation of the F-score directly within word vectors, thereby establishing a direct link to the assessment of sentence similarity. In experiments with widely-used pre-trained embeddings and benchmarks, we show that our subspace-based set operations consistently outperform vector-based ones in both sentence similarity and set retrieval tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2210_13034
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle Subspace Representations for Soft Set Operations and Sentence Similarities
Ishibashi, Yoichi
Yokoi, Sho
Sudoh, Katsuhito
Nakamura, Satoshi
Computation and Language
Machine Learning
In the field of natural language processing (NLP), continuous vector representations are crucial for capturing the semantic meanings of individual words. Yet, when it comes to the representations of sets of words, the conventional vector-based approaches often struggle with expressiveness and lack the essential set operations such as union, intersection, and complement. Inspired by quantum logic, we realize the representation of word sets and corresponding set operations within pre-trained word embedding spaces. By grounding our approach in the linear subspaces, we enable efficient computation of various set operations and facilitate the soft computation of membership functions within continuous spaces. Moreover, we allow for the computation of the F-score directly within word vectors, thereby establishing a direct link to the assessment of sentence similarity. In experiments with widely-used pre-trained embeddings and benchmarks, we show that our subspace-based set operations consistently outperform vector-based ones in both sentence similarity and set retrieval tasks.
title Subspace Representations for Soft Set Operations and Sentence Similarities
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2210.13034