SUNAR: Semantic Uncertainty based Neighborhood Aware Retrieval for Complex QA

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Venktesh, V, Rathee, Mandeep, Anand, Avishek
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912289532149760
author Venktesh, V
Rathee, Mandeep
Anand, Avishek
author_facet Venktesh, V
Rathee, Mandeep
Anand, Avishek
contents Complex question-answering (QA) systems face significant challenges in retrieving and reasoning over information that addresses multi-faceted queries. While large language models (LLMs) have advanced the reasoning capabilities of these systems, the bounded-recall problem persists, where procuring all relevant documents in first-stage retrieval remains a challenge. Missing pertinent documents at this stage leads to performance degradation that cannot be remedied in later stages, especially given the limited context windows of LLMs which necessitate high recall at smaller retrieval depths. In this paper, we introduce SUNAR, a novel approach that leverages LLMs to guide a Neighborhood Aware Retrieval process. SUNAR iteratively explores a neighborhood graph of documents, dynamically promoting or penalizing documents based on uncertainty estimates from interim LLM-generated answer candidates. We validate our approach through extensive experiments on two complex QA datasets. Our results show that SUNAR significantly outperforms existing retrieve-and-reason baselines, achieving up to a 31.84% improvement in performance over existing state-of-the-art methods for complex QA.
format Preprint
id arxiv_https___arxiv_org_abs_2503_17990
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SUNAR: Semantic Uncertainty based Neighborhood Aware Retrieval for Complex QA
Venktesh, V
Rathee, Mandeep
Anand, Avishek
Information Retrieval
Complex question-answering (QA) systems face significant challenges in retrieving and reasoning over information that addresses multi-faceted queries. While large language models (LLMs) have advanced the reasoning capabilities of these systems, the bounded-recall problem persists, where procuring all relevant documents in first-stage retrieval remains a challenge. Missing pertinent documents at this stage leads to performance degradation that cannot be remedied in later stages, especially given the limited context windows of LLMs which necessitate high recall at smaller retrieval depths. In this paper, we introduce SUNAR, a novel approach that leverages LLMs to guide a Neighborhood Aware Retrieval process. SUNAR iteratively explores a neighborhood graph of documents, dynamically promoting or penalizing documents based on uncertainty estimates from interim LLM-generated answer candidates. We validate our approach through extensive experiments on two complex QA datasets. Our results show that SUNAR significantly outperforms existing retrieve-and-reason baselines, achieving up to a 31.84% improvement in performance over existing state-of-the-art methods for complex QA.
title SUNAR: Semantic Uncertainty based Neighborhood Aware Retrieval for Complex QA
topic Information Retrieval
url https://arxiv.org/abs/2503.17990