Consensus or Conflict? Fine-Grained Evaluation of Conflicting Answers in Question-Answering
Fuente:
arXiv
Saved in:
| Main Authors: | Nachshoni, Eviatar, Cattan, Arie, Amar, Shmuel, Shapira, Ori, Dagan, Ido |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Unifying Scheme for Extractive Content Selection Tasks
by: Amar, Shmuel, et al.
Published: (2025)
by: Amar, Shmuel, et al.
Published: (2025)
EventFull: Complete and Consistent Event Relation Annotation
by: Eirew, Alon, et al.
Published: (2024)
by: Eirew, Alon, et al.
Published: (2024)
QAPyramid: Fine-grained Evaluation of Content Selection for Text Summarization
by: Zhang, Shiyue, et al.
Published: (2024)
by: Zhang, Shiyue, et al.
Published: (2024)
Multi-Review Fusion-in-Context
by: Slobodkin, Aviv, et al.
Published: (2024)
by: Slobodkin, Aviv, et al.
Published: (2024)
Attribute First, then Generate: Locally-attributable Grounded Text Generation
by: Slobodkin, Aviv, et al.
Published: (2024)
by: Slobodkin, Aviv, et al.
Published: (2024)
DRAGged into Conflicts: Detecting and Addressing Conflicting Sources in Search-Augmented LLMs
by: Cattan, Arie, et al.
Published: (2025)
by: Cattan, Arie, et al.
Published: (2025)
SEAM: A Stochastic Benchmark for Multi-Document Tasks
by: Lior, Gili, et al.
Published: (2024)
by: Lior, Gili, et al.
Published: (2024)
Explicating the Implicit: Argument Detection Beyond Sentence Boundaries
by: Roit, Paul, et al.
Published: (2024)
by: Roit, Paul, et al.
Published: (2024)
CLATTER: Comprehensive Entailment Reasoning for Hallucination Detection
by: Eliav, Ron, et al.
Published: (2025)
by: Eliav, Ron, et al.
Published: (2025)
The Power of Summary-Source Alignments
by: Ernst, Ori, et al.
Published: (2024)
by: Ernst, Ori, et al.
Published: (2024)
QA-Noun: Representing Nominal Semantics via Natural Language Question-Answer Pairs
by: Tseytlin, Maria, et al.
Published: (2025)
by: Tseytlin, Maria, et al.
Published: (2025)
Localizing Factual Inconsistencies in Attributable Text Generation
by: Cattan, Arie, et al.
Published: (2024)
by: Cattan, Arie, et al.
Published: (2024)
Open Domain Question Answering with Conflicting Contexts
by: Liu, Siyi, et al.
Published: (2024)
by: Liu, Siyi, et al.
Published: (2024)
Question Answering under Temporal Conflict: Evaluating and Organizing Evolving Knowledge with LLMs
by: Özer, Atahan, et al.
Published: (2025)
by: Özer, Atahan, et al.
Published: (2025)
Mitigating Knowledge Conflicts in Language Model-Driven Question Answering
by: Cao, Han, et al.
Published: (2024)
by: Cao, Han, et al.
Published: (2024)
Information Types in Product Reviews
by: Shapira, Ori, et al.
Published: (2025)
by: Shapira, Ori, et al.
Published: (2025)
When Evidence Conflicts: Uncertainty and Order Effects in Retrieval-Augmented Biomedical Question Answering
by: Han, Yikun, et al.
Published: (2026)
by: Han, Yikun, et al.
Published: (2026)
FREB-TQA: A Fine-Grained Robustness Evaluation Benchmark for Table Question Answering
by: Zhou, Wei, et al.
Published: (2024)
by: Zhou, Wei, et al.
Published: (2024)
Adaptive Question Answering: Enhancing Language Model Proficiency for Addressing Knowledge Conflicts with Source Citations
by: Shaier, Sagi, et al.
Published: (2024)
by: Shaier, Sagi, et al.
Published: (2024)
Evaluating Answer Reranking Strategies in Time-sensitive Question Answering
by: Kardan, Mehmet, et al.
Published: (2025)
by: Kardan, Mehmet, et al.
Published: (2025)
Effective QA-driven Annotation of Predicate-Argument Relations Across Languages
by: Davidov, Jonathan, et al.
Published: (2026)
by: Davidov, Jonathan, et al.
Published: (2026)
Narrowing the Knowledge Evaluation Gap: Open-Domain Question Answering with Multi-Granularity Answers
by: Yona, Gal, et al.
Published: (2024)
by: Yona, Gal, et al.
Published: (2024)
Beyond Pairwise: Global Zero-shot Temporal Graph Generation
by: Eirew, Alon, et al.
Published: (2025)
by: Eirew, Alon, et al.
Published: (2025)
Right Answer, Wrong Score: Uncovering the Inconsistencies of LLM Evaluation in Multiple-Choice Question Answering
by: Molfese, Francesco Maria, et al.
Published: (2025)
by: Molfese, Francesco Maria, et al.
Published: (2025)
CFMatch: Aligning Automated Answer Equivalence Evaluation with Expert Judgments For Open-Domain Question Answering
by: Li, Zongxia, et al.
Published: (2024)
by: Li, Zongxia, et al.
Published: (2024)
Question: How do Large Language Models perform on the Question Answering tasks? Answer:
by: Fischer, Kevin, et al.
Published: (2024)
by: Fischer, Kevin, et al.
Published: (2024)
Question Answering with LLMs and Learning from Answer Sets
by: Borroto, Manuel, et al.
Published: (2025)
by: Borroto, Manuel, et al.
Published: (2025)
Graph Guided Question Answer Generation for Procedural Question-Answering
by: Pham, Hai X., et al.
Published: (2024)
by: Pham, Hai X., et al.
Published: (2024)
GenerationPrograms: Fine-grained Attribution with Executable Programs
by: Wan, David, et al.
Published: (2025)
by: Wan, David, et al.
Published: (2025)
ConflictBench: Evaluating Human-AI Conflict via Interactive and Visually Grounded Environments
by: Zhao, Weixiang, et al.
Published: (2026)
by: Zhao, Weixiang, et al.
Published: (2026)
ConflictBank: A Benchmark for Evaluating the Influence of Knowledge Conflicts in LLM
by: Su, Zhaochen, et al.
Published: (2024)
by: Su, Zhaochen, et al.
Published: (2024)
Is That Your Final Answer? Test-Time Scaling Improves Selective Question Answering
by: Jurayj, William, et al.
Published: (2025)
by: Jurayj, William, et al.
Published: (2025)
How often do Answers Change? Estimating Recency Requirements in Question Answering
by: Piryani, Bhawna, et al.
Published: (2026)
by: Piryani, Bhawna, et al.
Published: (2026)
Reverse Question Answering: Can an LLM Write a Question so Hard (or Bad) that it Can't Answer?
by: Balepur, Nishant, et al.
Published: (2024)
by: Balepur, Nishant, et al.
Published: (2024)
General Table Question Answering via Answer-Formula Joint Generation
by: Wang, Zhongyuan, et al.
Published: (2025)
by: Wang, Zhongyuan, et al.
Published: (2025)
A Dataset of Open-Domain Question Answering with Multiple-Span Answers
by: Luo, Zhiyi, et al.
Published: (2024)
by: Luo, Zhiyi, et al.
Published: (2024)
Answering Questions by Meta-Reasoning over Multiple Chains of Thought
by: Yoran, Ori, et al.
Published: (2023)
by: Yoran, Ori, et al.
Published: (2023)
Measuring the Effect of Transcription Noise on Downstream Language Understanding Tasks
by: Shapira, Ori, et al.
Published: (2025)
by: Shapira, Ori, et al.
Published: (2025)
Answer is All You Need: Instruction-following Text Embedding via Answering the Question
by: Peng, Letian, et al.
Published: (2024)
by: Peng, Letian, et al.
Published: (2024)
Dont Add, dont Miss: Effective Content Preserving Generation from Pre-Selected Text Spans
by: Slobodkin, Aviv, et al.
Published: (2023)
by: Slobodkin, Aviv, et al.
Published: (2023)
Similar Items
-
A Unifying Scheme for Extractive Content Selection Tasks
by: Amar, Shmuel, et al.
Published: (2025) -
EventFull: Complete and Consistent Event Relation Annotation
by: Eirew, Alon, et al.
Published: (2024) -
QAPyramid: Fine-grained Evaluation of Content Selection for Text Summarization
by: Zhang, Shiyue, et al.
Published: (2024) -
Multi-Review Fusion-in-Context
by: Slobodkin, Aviv, et al.
Published: (2024) -
Attribute First, then Generate: Locally-attributable Grounded Text Generation
by: Slobodkin, Aviv, et al.
Published: (2024)