Consensus or Conflict? Fine-Grained Evaluation of Conflicting Answers in Question-Answering

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nachshoni, Eviatar, Cattan, Arie, Amar, Shmuel, Shapira, Ori, Dagan, Ido
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915448743788544
author Nachshoni, Eviatar
Cattan, Arie
Amar, Shmuel
Shapira, Ori
Dagan, Ido
author_facet Nachshoni, Eviatar
Cattan, Arie
Amar, Shmuel
Shapira, Ori
Dagan, Ido
contents Large Language Models (LLMs) have demonstrated strong performance in question answering (QA) tasks. However, Multi-Answer Question Answering (MAQA), where a question may have several valid answers, remains challenging. Traditional QA settings often assume consistency across evidences, but MAQA can involve conflicting answers. Constructing datasets that reflect such conflicts is costly and labor-intensive, while existing benchmarks often rely on synthetic data, restrict the task to yes/no questions, or apply unverified automated annotation. To advance research in this area, we extend the conflict-aware MAQA setting to require models not only to identify all valid answers, but also to detect specific conflicting answer pairs, if any. To support this task, we introduce a novel cost-effective methodology for leveraging fact-checking datasets to construct NATCONFQA, a new benchmark for realistic, conflict-aware MAQA, enriched with detailed conflict labels, for all answer pairs. We evaluate eight high-end LLMs on NATCONFQA, revealing their fragility in handling various types of conflicts and the flawed strategies they employ to resolve them.
format Preprint
id arxiv_https___arxiv_org_abs_2508_12355
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Consensus or Conflict? Fine-Grained Evaluation of Conflicting Answers in Question-Answering
Nachshoni, Eviatar
Cattan, Arie
Amar, Shmuel
Shapira, Ori
Dagan, Ido
Computation and Language
Large Language Models (LLMs) have demonstrated strong performance in question answering (QA) tasks. However, Multi-Answer Question Answering (MAQA), where a question may have several valid answers, remains challenging. Traditional QA settings often assume consistency across evidences, but MAQA can involve conflicting answers. Constructing datasets that reflect such conflicts is costly and labor-intensive, while existing benchmarks often rely on synthetic data, restrict the task to yes/no questions, or apply unverified automated annotation. To advance research in this area, we extend the conflict-aware MAQA setting to require models not only to identify all valid answers, but also to detect specific conflicting answer pairs, if any. To support this task, we introduce a novel cost-effective methodology for leveraging fact-checking datasets to construct NATCONFQA, a new benchmark for realistic, conflict-aware MAQA, enriched with detailed conflict labels, for all answer pairs. We evaluate eight high-end LLMs on NATCONFQA, revealing their fragility in handling various types of conflicts and the flawed strategies they employ to resolve them.
title Consensus or Conflict? Fine-Grained Evaluation of Conflicting Answers in Question-Answering
topic Computation and Language
url https://arxiv.org/abs/2508.12355