Soft Similarity and Soft Cosine Measure: Similarity of Features in Vector Space Model

Fuente: Redalyc
Gespeichert in:
Bibliographische Detailangaben
1. Verfasser: Grigori Sidorov
Format: Artículo científico
Sprache:en
Veröffentlicht: Instituto Politécnico Nacional 2014
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1876445721200164864
author Grigori Sidorov
author_facet Grigori Sidorov
contents Soft Similarity and Soft Cosine Measure: Similarity of Features in Vector Space Model Grigori Sidorov Alexander Gelbukh Helena Gómez-Adorno David Pinto Computación grams syntactic n Soft similarity vector space model soft cosine measure We show how to consider similarity between features for calculation of similarity of objects in the Vector Space Model (VSM) for machine learning algorithms and other classes of methods that involve similarity between objects. Unlike LSA, we assume that similarity between features is known (say, from a synonym dictionary) and does not need to be learned from the data. We call the proposed similarity measure soft similarity. Similarity between features is common, for example, in natural language processing: words, n-grams, or syntactic n-grams can be somewhat different (which makes them different features) but still have much in common: for example, words “play” and “game” are different but related. When there is no similarity between features then our soft similarity measure is equal to the standard similarity. For this, we generalize the well-known cosine similarity measure in VSM by introducing what we call “soft cosine measure”. We propose various formulas for exact or approximate calculation of the soft cosine measure. For example, in one of them we consider for VSM a new feature space consisting of pairs of the original features weighted by their similarity. Again, for features that bear no similarity to each other, our formulas reduce to the standard cosine measure. Our experiments show that our soft cosine measure provides better performance in our case study: entrance exams question answering task at CLEF. In these experiments, we use syntactic n-grams as features and Levenshtein distance as the similarity between n-grams, measured either in characters or in elements of n-grams. 2014 artículo científico 1405-5546 https://www.redalyc.org/articulo.oa?id=61532067007 en http://www.redalyc.org/revista.oa?id=615 Computación y Sistemas application/pdf Instituto Politécnico Nacional Computación y Sistemas (México) Num.3 Vol.18
format Artículo científico
id redalyc_61532067007
institution Redalyc
language en
publishDate 2014
publisher Instituto Politécnico Nacional
spellingShingle Soft Similarity and Soft Cosine Measure: Similarity of Features in Vector Space Model
Grigori Sidorov
Computación
grams
syntactic n
Soft similarity
vector space model
soft cosine measure
Soft Similarity and Soft Cosine Measure: Similarity of Features in Vector Space Model Grigori Sidorov Alexander Gelbukh Helena Gómez-Adorno David Pinto Computación grams syntactic n Soft similarity vector space model soft cosine measure We show how to consider similarity between features for calculation of similarity of objects in the Vector Space Model (VSM) for machine learning algorithms and other classes of methods that involve similarity between objects. Unlike LSA, we assume that similarity between features is known (say, from a synonym dictionary) and does not need to be learned from the data. We call the proposed similarity measure soft similarity. Similarity between features is common, for example, in natural language processing: words, n-grams, or syntactic n-grams can be somewhat different (which makes them different features) but still have much in common: for example, words “play” and “game” are different but related. When there is no similarity between features then our soft similarity measure is equal to the standard similarity. For this, we generalize the well-known cosine similarity measure in VSM by introducing what we call “soft cosine measure”. We propose various formulas for exact or approximate calculation of the soft cosine measure. For example, in one of them we consider for VSM a new feature space consisting of pairs of the original features weighted by their similarity. Again, for features that bear no similarity to each other, our formulas reduce to the standard cosine measure. Our experiments show that our soft cosine measure provides better performance in our case study: entrance exams question answering task at CLEF. In these experiments, we use syntactic n-grams as features and Levenshtein distance as the similarity between n-grams, measured either in characters or in elements of n-grams. 2014 artículo científico 1405-5546 https://www.redalyc.org/articulo.oa?id=61532067007 en http://www.redalyc.org/revista.oa?id=615 Computación y Sistemas application/pdf Instituto Politécnico Nacional Computación y Sistemas (México) Num.3 Vol.18
title Soft Similarity and Soft Cosine Measure: Similarity of Features in Vector Space Model
topic Computación
grams
syntactic n
Soft similarity
vector space model
soft cosine measure
url https://www.redalyc.org/articulo.oa?id=61532067007