Answering Multimodal Exclusion Queries with Lightweight Sparse Disentangled Representations

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: J, Prachi, Bhatia, Sumit, Bedathur, Srikanta
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911020010700800
author J, Prachi
Bhatia, Sumit
Bedathur, Srikanta
author_facet J, Prachi
Bhatia, Sumit
Bedathur, Srikanta
contents Multimodal representations that enable cross-modal retrieval are widely used. However, these often lack interpretability making it difficult to explain the retrieved results. Solutions such as learning sparse disentangled representations are typically guided by the text tokens in the data, making the dimensionality of the resulting embeddings very high. We propose an approach that generates smaller dimensionality fixed-size embeddings that are not only disentangled but also offer better control for retrieval tasks. We demonstrate their utility using challenging exclusion queries over MSCOCO and Conceptual Captions benchmarks. Our experiments show that our approach is superior to traditional dense models such as CLIP, BLIP and VISTA (gains up to 11% in AP@10), as well as sparse disentangled models like VDR (gains up to 21% in AP@10). We also present qualitative results to further underline the interpretability of disentangled representations.
format Preprint
id arxiv_https___arxiv_org_abs_2504_03184
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Answering Multimodal Exclusion Queries with Lightweight Sparse Disentangled Representations
J, Prachi
Bhatia, Sumit
Bedathur, Srikanta
Information Retrieval
Multimodal representations that enable cross-modal retrieval are widely used. However, these often lack interpretability making it difficult to explain the retrieved results. Solutions such as learning sparse disentangled representations are typically guided by the text tokens in the data, making the dimensionality of the resulting embeddings very high. We propose an approach that generates smaller dimensionality fixed-size embeddings that are not only disentangled but also offer better control for retrieval tasks. We demonstrate their utility using challenging exclusion queries over MSCOCO and Conceptual Captions benchmarks. Our experiments show that our approach is superior to traditional dense models such as CLIP, BLIP and VISTA (gains up to 11% in AP@10), as well as sparse disentangled models like VDR (gains up to 21% in AP@10). We also present qualitative results to further underline the interpretability of disentangled representations.
title Answering Multimodal Exclusion Queries with Lightweight Sparse Disentangled Representations
topic Information Retrieval
url https://arxiv.org/abs/2504.03184