Mind the Gap: A Generalized Approach for Cross-Modal Embedding Alignment

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yadav, Arihan, McMillan, Alan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929569478475776
author Yadav, Arihan
McMillan, Alan
author_facet Yadav, Arihan
McMillan, Alan
contents Retrieval-Augmented Generation (RAG) systems enhance text generation by incorporating external knowledge but often struggle when retrieving context across different text modalities due to semantic gaps. We introduce a generalized projection-based method, inspired by adapter modules in transfer learning, that efficiently bridges these gaps between various text types, such as programming code and pseudocode, or English and French sentences. Our approach emphasizes speed, accuracy, and data efficiency, requiring minimal resources for training and inference. By aligning embeddings from heterogeneous text modalities into a unified space through a lightweight projection network, our model significantly outperforms traditional retrieval methods like the Okapi BM25 algorithm and models like Dense Passage Retrieval (DPR), while approaching the accuracy of Sentence Transformers. Extensive evaluations demonstrate the effectiveness and generalizability of our method across different tasks, highlighting its potential for real-time, resource-constrained applications.
format Preprint
id arxiv_https___arxiv_org_abs_2410_23437
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Mind the Gap: A Generalized Approach for Cross-Modal Embedding Alignment
Yadav, Arihan
McMillan, Alan
Machine Learning
Computation and Language
Information Retrieval
H.3.3; I.2.7; I.2.6
Retrieval-Augmented Generation (RAG) systems enhance text generation by incorporating external knowledge but often struggle when retrieving context across different text modalities due to semantic gaps. We introduce a generalized projection-based method, inspired by adapter modules in transfer learning, that efficiently bridges these gaps between various text types, such as programming code and pseudocode, or English and French sentences. Our approach emphasizes speed, accuracy, and data efficiency, requiring minimal resources for training and inference. By aligning embeddings from heterogeneous text modalities into a unified space through a lightweight projection network, our model significantly outperforms traditional retrieval methods like the Okapi BM25 algorithm and models like Dense Passage Retrieval (DPR), while approaching the accuracy of Sentence Transformers. Extensive evaluations demonstrate the effectiveness and generalizability of our method across different tasks, highlighting its potential for real-time, resource-constrained applications.
title Mind the Gap: A Generalized Approach for Cross-Modal Embedding Alignment
topic Machine Learning
Computation and Language
Information Retrieval
H.3.3; I.2.7; I.2.6
url https://arxiv.org/abs/2410.23437