Saved in:
Bibliographic Details
Main Authors: Liu, Tian Yu, Soatto, Stefano
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2410.16431
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918459774861312
author Liu, Tian Yu
Soatto, Stefano
author_facet Liu, Tian Yu
Soatto, Stefano
contents The semantic similarity between sample expressions measures the distance between their latent 'meaning'. These meanings are themselves typically represented by textual expressions. We propose a novel approach whereby the semantic similarity among textual expressions is based not on other expressions they can be rephrased as, but rather based on the imagery they evoke. While this is not possible with humans, generative models allow us to easily visualize and compare generated images, or their distribution, evoked by a textual prompt. Therefore, we characterize the semantic similarity between two textual expressions simply as the distance between image distributions they induce, or 'conjure.' We show that by choosing the Jeffreys divergence between the reverse-time diffusion stochastic differential equations (SDEs) induced by each textual expression, this can be directly computed via Monte-Carlo sampling. Our method contributes a novel perspective on semantic similarity that not only aligns with human-annotated scores, but also opens up new avenues for the evaluation of text-conditioned generative models while offering better interpretability of their learnt representations.
format Preprint
id arxiv_https___arxiv_org_abs_2410_16431
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Conjuring Semantic Similarity
Liu, Tian Yu
Soatto, Stefano
Artificial Intelligence
The semantic similarity between sample expressions measures the distance between their latent 'meaning'. These meanings are themselves typically represented by textual expressions. We propose a novel approach whereby the semantic similarity among textual expressions is based not on other expressions they can be rephrased as, but rather based on the imagery they evoke. While this is not possible with humans, generative models allow us to easily visualize and compare generated images, or their distribution, evoked by a textual prompt. Therefore, we characterize the semantic similarity between two textual expressions simply as the distance between image distributions they induce, or 'conjure.' We show that by choosing the Jeffreys divergence between the reverse-time diffusion stochastic differential equations (SDEs) induced by each textual expression, this can be directly computed via Monte-Carlo sampling. Our method contributes a novel perspective on semantic similarity that not only aligns with human-annotated scores, but also opens up new avenues for the evaluation of text-conditioned generative models while offering better interpretability of their learnt representations.
title Conjuring Semantic Similarity
topic Artificial Intelligence
url https://arxiv.org/abs/2410.16431