Saved in:
Bibliographic Details
Main Authors: Lee, Youngwon, Hwang, Seung-won, Wu, Ruofan, Yan, Feng, Xu, Danmei, Akkad, Moutasem, Yao, Zhewei, He, Yuxiong
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2502.10352
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915182256586752
author Lee, Youngwon
Hwang, Seung-won
Wu, Ruofan
Yan, Feng
Xu, Danmei
Akkad, Moutasem
Yao, Zhewei
He, Yuxiong
author_facet Lee, Youngwon
Hwang, Seung-won
Wu, Ruofan
Yan, Feng
Xu, Danmei
Akkad, Moutasem
Yao, Zhewei
He, Yuxiong
contents In this work, we tackle the challenge of disambiguating queries in retrieval-augmented generation (RAG) to diverse yet answerable interpretations. State-of-the-arts follow a Diversify-then-Verify (DtV) pipeline, where diverse interpretations are generated by an LLM, later used as search queries to retrieve supporting passages. Such a process may introduce noise in either interpretations or retrieval, particularly in enterprise settings, where LLMs -- trained on static data -- may struggle with domain-specific disambiguations. Thus, a post-hoc verification phase is introduced to prune noises. Our distinction is to unify diversification with verification by incorporating feedback from retriever and generator early on. This joint approach improves both efficiency and robustness by reducing reliance on multiple retrieval and inference steps, which are susceptible to cascading errors. We validate the efficiency and effectiveness of our method, Verified-Diversification with Consolidation (VERDICT), on the widely adopted ASQA benchmark to achieve diverse yet verifiable interpretations. Empirical results show that VERDICT improves grounding-aware F1 score by an average of 23% over the strongest baseline across different backbone LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2502_10352
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Agentic Verification for Ambiguous Query Disambiguation
Lee, Youngwon
Hwang, Seung-won
Wu, Ruofan
Yan, Feng
Xu, Danmei
Akkad, Moutasem
Yao, Zhewei
He, Yuxiong
Computation and Language
In this work, we tackle the challenge of disambiguating queries in retrieval-augmented generation (RAG) to diverse yet answerable interpretations. State-of-the-arts follow a Diversify-then-Verify (DtV) pipeline, where diverse interpretations are generated by an LLM, later used as search queries to retrieve supporting passages. Such a process may introduce noise in either interpretations or retrieval, particularly in enterprise settings, where LLMs -- trained on static data -- may struggle with domain-specific disambiguations. Thus, a post-hoc verification phase is introduced to prune noises. Our distinction is to unify diversification with verification by incorporating feedback from retriever and generator early on. This joint approach improves both efficiency and robustness by reducing reliance on multiple retrieval and inference steps, which are susceptible to cascading errors. We validate the efficiency and effectiveness of our method, Verified-Diversification with Consolidation (VERDICT), on the widely adopted ASQA benchmark to achieve diverse yet verifiable interpretations. Empirical results show that VERDICT improves grounding-aware F1 score by an average of 23% over the strongest baseline across different backbone LLMs.
title Agentic Verification for Ambiguous Query Disambiguation
topic Computation and Language
url https://arxiv.org/abs/2502.10352