QuOTE: Question-Oriented Text Embeddings

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Neeser, Andrew, Latimer, Kaylen, Khatri, Aadyant, Latimer, Chris, Ramakrishnan, Naren
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917924666605568
author Neeser, Andrew
Latimer, Kaylen
Khatri, Aadyant
Latimer, Chris
Ramakrishnan, Naren
author_facet Neeser, Andrew
Latimer, Kaylen
Khatri, Aadyant
Latimer, Chris
Ramakrishnan, Naren
contents We present QuOTE (Question-Oriented Text Embeddings), a novel enhancement to retrieval-augmented generation (RAG) systems, aimed at improving document representation for accurate and nuanced retrieval. Unlike traditional RAG pipelines, which rely on embedding raw text chunks, QuOTE augments chunks with hypothetical questions that the chunk can potentially answer, enriching the representation space. This better aligns document embeddings with user query semantics, and helps address issues such as ambiguity and context-dependent relevance. Through extensive experiments across diverse benchmarks, we demonstrate that QuOTE significantly enhances retrieval accuracy, including in multi-hop question-answering tasks. Our findings highlight the versatility of question generation as a fundamental indexing strategy, opening new avenues for integrating question generation into retrieval-based AI pipelines.
format Preprint
id arxiv_https___arxiv_org_abs_2502_10976
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle QuOTE: Question-Oriented Text Embeddings
Neeser, Andrew
Latimer, Kaylen
Khatri, Aadyant
Latimer, Chris
Ramakrishnan, Naren
Information Retrieval
Artificial Intelligence
Computation and Language
Machine Learning
H.3
We present QuOTE (Question-Oriented Text Embeddings), a novel enhancement to retrieval-augmented generation (RAG) systems, aimed at improving document representation for accurate and nuanced retrieval. Unlike traditional RAG pipelines, which rely on embedding raw text chunks, QuOTE augments chunks with hypothetical questions that the chunk can potentially answer, enriching the representation space. This better aligns document embeddings with user query semantics, and helps address issues such as ambiguity and context-dependent relevance. Through extensive experiments across diverse benchmarks, we demonstrate that QuOTE significantly enhances retrieval accuracy, including in multi-hop question-answering tasks. Our findings highlight the versatility of question generation as a fundamental indexing strategy, opening new avenues for integrating question generation into retrieval-based AI pipelines.
title QuOTE: Question-Oriented Text Embeddings
topic Information Retrieval
Artificial Intelligence
Computation and Language
Machine Learning
H.3
url https://arxiv.org/abs/2502.10976