Narrowing the Knowledge Evaluation Gap: Open-Domain Question Answering with Multi-Granularity Answers
Fuente:
arXiv
Salvato in:
| Autori principali: | Yona, Gal, Aharoni, Roee, Geva, Mor |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Can Large Language Models Faithfully Express Their Intrinsic Uncertainty in Words?
di: Yona, Gal, et al.
Pubblicazione: (2024)
di: Yona, Gal, et al.
Pubblicazione: (2024)
Hallucinations Undermine Trust; Metacognition is a Way Forward
di: Yona, Gal, et al.
Pubblicazione: (2026)
di: Yona, Gal, et al.
Pubblicazione: (2026)
Keep Guessing? When Considering Inference Scaling, Mind the Baselines
di: Yona, Gal, et al.
Pubblicazione: (2024)
di: Yona, Gal, et al.
Pubblicazione: (2024)
Thinking to Recall: How Reasoning Unlocks Parametric Knowledge in LLMs
di: Gekhman, Zorik, et al.
Pubblicazione: (2026)
di: Gekhman, Zorik, et al.
Pubblicazione: (2026)
Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?
di: Gekhman, Zorik, et al.
Pubblicazione: (2024)
di: Gekhman, Zorik, et al.
Pubblicazione: (2024)
Friends and Grandmothers in Silico: Localizing Entity Cells in Language Models
di: Yona, Itay, et al.
Pubblicazione: (2026)
di: Yona, Itay, et al.
Pubblicazione: (2026)
Performance Gap in Entity Knowledge Extraction Across Modalities in Vision Language Models
di: Cohen, Ido, et al.
Pubblicazione: (2024)
di: Cohen, Ido, et al.
Pubblicazione: (2024)
Estimating Knowledge in Large Language Models Without Generating a Single Token
di: Gottesman, Daniela, et al.
Pubblicazione: (2024)
di: Gottesman, Daniela, et al.
Pubblicazione: (2024)
A Chain-of-Thought Is as Strong as Its Weakest Link: A Benchmark for Verifiers of Reasoning Chains
di: Jacovi, Alon, et al.
Pubblicazione: (2024)
di: Jacovi, Alon, et al.
Pubblicazione: (2024)
CFMatch: Aligning Automated Answer Equivalence Evaluation with Expert Judgments For Open-Domain Question Answering
di: Li, Zongxia, et al.
Pubblicazione: (2024)
di: Li, Zongxia, et al.
Pubblicazione: (2024)
A Dataset of Open-Domain Question Answering with Multiple-Span Answers
di: Luo, Zhiyi, et al.
Pubblicazione: (2024)
di: Luo, Zhiyi, et al.
Pubblicazione: (2024)
Inferring Functionality of Attention Heads from their Parameters
di: Elhelo, Amit, et al.
Pubblicazione: (2024)
di: Elhelo, Amit, et al.
Pubblicazione: (2024)
Dr3: Ask Large Language Models Not to Give Off-Topic Answers in Open Domain Multi-Hop Question Answering
di: Gao, Yuan, et al.
Pubblicazione: (2024)
di: Gao, Yuan, et al.
Pubblicazione: (2024)
Rethinking Selective Knowledge Distillation
di: Tavor, Almog, et al.
Pubblicazione: (2026)
di: Tavor, Almog, et al.
Pubblicazione: (2026)
Preventing Rogue Agents Improves Multi-Agent Collaboration
di: Barbi, Ohav, et al.
Pubblicazione: (2025)
di: Barbi, Ohav, et al.
Pubblicazione: (2025)
Faithfulness Metrics Don't Measure Faithfulness: A Meta-Evaluation with Ground Truth
di: Gur-Arieh, Yoav, et al.
Pubblicazione: (2026)
di: Gur-Arieh, Yoav, et al.
Pubblicazione: (2026)
Beneath the Surface of Consistency: Exploring Cross-lingual Knowledge Representation Sharing in LLMs
di: Ifergan, Maxim, et al.
Pubblicazione: (2024)
di: Ifergan, Maxim, et al.
Pubblicazione: (2024)
Consensus or Conflict? Fine-Grained Evaluation of Conflicting Answers in Question-Answering
di: Nachshoni, Eviatar, et al.
Pubblicazione: (2025)
di: Nachshoni, Eviatar, et al.
Pubblicazione: (2025)
Evaluating Answer Reranking Strategies in Time-sensitive Question Answering
di: Kardan, Mehmet, et al.
Pubblicazione: (2025)
di: Kardan, Mehmet, et al.
Pubblicazione: (2025)
Few-Shot Data Synthesis for Open Domain Multi-Hop Question Answering
di: Chen, Mingda, et al.
Pubblicazione: (2023)
di: Chen, Mingda, et al.
Pubblicazione: (2023)
Answerability in Retrieval-Augmented Open-Domain Question Answering
di: Abdumalikov, Rustam, et al.
Pubblicazione: (2024)
di: Abdumalikov, Rustam, et al.
Pubblicazione: (2024)
Open Domain Question Answering with Conflicting Contexts
di: Liu, Siyi, et al.
Pubblicazione: (2024)
di: Liu, Siyi, et al.
Pubblicazione: (2024)
mFACE: Multilingual Summarization with Factual Consistency Evaluation
di: Aharoni, Roee, et al.
Pubblicazione: (2022)
di: Aharoni, Roee, et al.
Pubblicazione: (2022)
Constructing Interpretable Features from Compositional Neuron Groups
di: Shafran, Or, et al.
Pubblicazione: (2025)
di: Shafran, Or, et al.
Pubblicazione: (2025)
Eliciting Textual Descriptions from Representations of Continuous Prompts
di: Ramati, Dana, et al.
Pubblicazione: (2024)
di: Ramati, Dana, et al.
Pubblicazione: (2024)
KazQAD: Kazakh Open-Domain Question Answering Dataset
di: Yeshpanov, Rustem, et al.
Pubblicazione: (2024)
di: Yeshpanov, Rustem, et al.
Pubblicazione: (2024)
Generator-Retriever-Generator Approach for Open-Domain Question Answering
di: Abdallah, Abdelrahman, et al.
Pubblicazione: (2023)
di: Abdallah, Abdelrahman, et al.
Pubblicazione: (2023)
QPaug: Question and Passage Augmentation for Open-Domain Question Answering of LLMs
di: Kim, Minsang, et al.
Pubblicazione: (2024)
di: Kim, Minsang, et al.
Pubblicazione: (2024)
Improving Retrieval Augmented Open-Domain Question-Answering with Vectorized Contexts
di: Chen, Zhuo, et al.
Pubblicazione: (2024)
di: Chen, Zhuo, et al.
Pubblicazione: (2024)
DebateQA: Evaluating Question Answering on Debatable Knowledge
di: Xu, Rongwu, et al.
Pubblicazione: (2024)
di: Xu, Rongwu, et al.
Pubblicazione: (2024)
Intrinsic Test of Unlearning Using Parametric Knowledge Traces
di: Hong, Yihuai, et al.
Pubblicazione: (2024)
di: Hong, Yihuai, et al.
Pubblicazione: (2024)
Cross-Granularity Hypergraph Retrieval-Augmented Generation for Multi-hop Question Answering
di: Wang, Changjian, et al.
Pubblicazione: (2025)
di: Wang, Changjian, et al.
Pubblicazione: (2025)
Right Answer, Wrong Score: Uncovering the Inconsistencies of LLM Evaluation in Multiple-Choice Question Answering
di: Molfese, Francesco Maria, et al.
Pubblicazione: (2025)
di: Molfese, Francesco Maria, et al.
Pubblicazione: (2025)
Routers Learn the Geometry of Their Experts: Geometric Coupling in Sparse Mixture-of-Experts
di: Ahrac, Sagi, et al.
Pubblicazione: (2026)
di: Ahrac, Sagi, et al.
Pubblicazione: (2026)
Mixing Mechanisms: How Language Models Retrieve Bound Entities In-Context
di: Gur-Arieh, Yoav, et al.
Pubblicazione: (2025)
di: Gur-Arieh, Yoav, et al.
Pubblicazione: (2025)
Disentangling MLP Neuron Weights in Vocabulary Space
di: Avrahamy, Asaf, et al.
Pubblicazione: (2026)
di: Avrahamy, Asaf, et al.
Pubblicazione: (2026)
Detecting (Un)answerability in Large Language Models with Linear Directions
di: Lavi, Maor Juliet, et al.
Pubblicazione: (2025)
di: Lavi, Maor Juliet, et al.
Pubblicazione: (2025)
Question: How do Large Language Models perform on the Question Answering tasks? Answer:
di: Fischer, Kevin, et al.
Pubblicazione: (2024)
di: Fischer, Kevin, et al.
Pubblicazione: (2024)
Denoising Table-Text Retrieval for Open-Domain Question Answering
di: Kang, Deokhyung, et al.
Pubblicazione: (2024)
di: Kang, Deokhyung, et al.
Pubblicazione: (2024)
Exploring Hint Generation Approaches in Open-Domain Question Answering
di: Mozafari, Jamshid, et al.
Pubblicazione: (2024)
di: Mozafari, Jamshid, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Can Large Language Models Faithfully Express Their Intrinsic Uncertainty in Words?
di: Yona, Gal, et al.
Pubblicazione: (2024) -
Hallucinations Undermine Trust; Metacognition is a Way Forward
di: Yona, Gal, et al.
Pubblicazione: (2026) -
Keep Guessing? When Considering Inference Scaling, Mind the Baselines
di: Yona, Gal, et al.
Pubblicazione: (2024) -
Thinking to Recall: How Reasoning Unlocks Parametric Knowledge in LLMs
di: Gekhman, Zorik, et al.
Pubblicazione: (2026) -
Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?
di: Gekhman, Zorik, et al.
Pubblicazione: (2024)