CondAmbigQA: A Benchmark and Dataset for Conditional Ambiguous Question Answering

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Li, Zongxi, Li, Yang, Xie, Haoran, Qin, S. Joe
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866916945348001792
author Li, Zongxi
Li, Yang
Xie, Haoran
Qin, S. Joe
author_facet Li, Zongxi
Li, Yang
Xie, Haoran
Qin, S. Joe
contents Users often assume that large language models (LLMs) share their cognitive alignment of context and intent, leading them to omit critical information in question-answering (QA) and produce ambiguous queries. Responses based on misaligned assumptions may be perceived as hallucinations. Therefore, identifying possible implicit assumptions is crucial in QA. To address this fundamental challenge, we propose Conditional Ambiguous Question-Answering (CondAmbigQA), a benchmark comprising 2,000 ambiguous queries and condition-aware evaluation metrics. Our study pioneers "conditions" as explicit contextual constraints that resolve ambiguities in QA tasks through retrieval-based annotation, where retrieved Wikipedia fragments help identify possible interpretations for a given query and annotate answers accordingly. Experiments demonstrate that models considering conditions before answering improve answer accuracy by 11.75%, with an additional 7.15% gain when conditions are explicitly provided. These results highlight that apparent hallucinations may stem from inherent query ambiguity rather than model failure, and demonstrate the effectiveness of condition reasoning in QA, providing researchers with tools for rigorous evaluation.
format Preprint
id arxiv_https___arxiv_org_abs_2502_01523
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CondAmbigQA: A Benchmark and Dataset for Conditional Ambiguous Question Answering
Li, Zongxi
Li, Yang
Xie, Haoran
Qin, S. Joe
Computation and Language
Users often assume that large language models (LLMs) share their cognitive alignment of context and intent, leading them to omit critical information in question-answering (QA) and produce ambiguous queries. Responses based on misaligned assumptions may be perceived as hallucinations. Therefore, identifying possible implicit assumptions is crucial in QA. To address this fundamental challenge, we propose Conditional Ambiguous Question-Answering (CondAmbigQA), a benchmark comprising 2,000 ambiguous queries and condition-aware evaluation metrics. Our study pioneers "conditions" as explicit contextual constraints that resolve ambiguities in QA tasks through retrieval-based annotation, where retrieved Wikipedia fragments help identify possible interpretations for a given query and annotate answers accordingly. Experiments demonstrate that models considering conditions before answering improve answer accuracy by 11.75%, with an additional 7.15% gain when conditions are explicitly provided. These results highlight that apparent hallucinations may stem from inherent query ambiguity rather than model failure, and demonstrate the effectiveness of condition reasoning in QA, providing researchers with tools for rigorous evaluation.
title CondAmbigQA: A Benchmark and Dataset for Conditional Ambiguous Question Answering
topic Computation and Language
url https://arxiv.org/abs/2502.01523