Description-Based Text Similarity

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ravfogel, Shauli, Pyatkin, Valentina, Cohen, Amir DN, Manevich, Avshalom, Goldberg, Yoav
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914884063592448
author Ravfogel, Shauli
Pyatkin, Valentina
Cohen, Amir DN
Manevich, Avshalom
Goldberg, Yoav
author_facet Ravfogel, Shauli
Pyatkin, Valentina
Cohen, Amir DN
Manevich, Avshalom
Goldberg, Yoav
contents Identifying texts with a given semantics is central for many information seeking scenarios. Similarity search over vector embeddings appear to be central to this ability, yet the similarity reflected in current text embeddings is corpus-driven, and is inconsistent and sub-optimal for many use cases. What, then, is a good notion of similarity for effective retrieval of text? We identify the need to search for texts based on abstract descriptions of their content, and the corresponding notion of \emph{description based similarity}. We demonstrate the inadequacy of current text embeddings and propose an alternative model that significantly improves when used in standard nearest neighbor search. The model is trained using positive and negative pairs sourced through prompting a LLM, demonstrating how data from LLMs can be used for creating new capabilities not immediately possible using the original model.
format Preprint
id arxiv_https___arxiv_org_abs_2305_12517
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Description-Based Text Similarity
Ravfogel, Shauli
Pyatkin, Valentina
Cohen, Amir DN
Manevich, Avshalom
Goldberg, Yoav
Computation and Language
Information Retrieval
Machine Learning
Identifying texts with a given semantics is central for many information seeking scenarios. Similarity search over vector embeddings appear to be central to this ability, yet the similarity reflected in current text embeddings is corpus-driven, and is inconsistent and sub-optimal for many use cases. What, then, is a good notion of similarity for effective retrieval of text? We identify the need to search for texts based on abstract descriptions of their content, and the corresponding notion of \emph{description based similarity}. We demonstrate the inadequacy of current text embeddings and propose an alternative model that significantly improves when used in standard nearest neighbor search. The model is trained using positive and negative pairs sourced through prompting a LLM, demonstrating how data from LLMs can be used for creating new capabilities not immediately possible using the original model.
title Description-Based Text Similarity
topic Computation and Language
Information Retrieval
Machine Learning
url https://arxiv.org/abs/2305.12517