How Reliable are LLMs as Knowledge Bases? Re-thinking Facutality and Consistency
Fuente:
arXiv
Saved in:
| Main Authors: | Zheng, Danna, Lapata, Mirella, Pan, Jeff Z. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Long-Form Information Alignment Evaluation Beyond Atomic Facts
by: Zheng, Danna, et al.
Published: (2025)
by: Zheng, Danna, et al.
Published: (2025)
Archer: A Human-Labeled Text-to-SQL Dataset with Arithmetic, Commonsense and Hypothetical Reasoning
by: Zheng, Danna, et al.
Published: (2024)
by: Zheng, Danna, et al.
Published: (2024)
TrustScore: Reference-Free Evaluation of LLM Response Trustworthiness
by: Zheng, Danna, et al.
Published: (2024)
by: Zheng, Danna, et al.
Published: (2024)
Reasoning about Intent for Ambiguous Requests
by: Saparina, Irina, et al.
Published: (2025)
by: Saparina, Irina, et al.
Published: (2025)
Debating for Better Reasoning: An Unsupervised Multimodal Approach
by: Adhikari, Ashutosh, et al.
Published: (2025)
by: Adhikari, Ashutosh, et al.
Published: (2025)
Disambiguate First, Parse Later: Generating Interpretations for Ambiguity Resolution in Semantic Parsing
by: Saparina, Irina, et al.
Published: (2025)
by: Saparina, Irina, et al.
Published: (2025)
Generating Visual Stories with Grounded and Coreferent Characters
by: Liu, Danyang, et al.
Published: (2024)
by: Liu, Danyang, et al.
Published: (2024)
Long-Context Reasoning Through Proxy-Based Chain-of-Thought Tuning
by: Li, Miao, et al.
Published: (2026)
by: Li, Miao, et al.
Published: (2026)
How Reliable are LLMs for Reasoning on the Re-ranking task?
by: Islam, Nafis Tanveer, et al.
Published: (2025)
by: Islam, Nafis Tanveer, et al.
Published: (2025)
Masking in Multi-hop QA: An Analysis of How Language Models Perform with Context Permutation
by: Huang, Wenyu, et al.
Published: (2025)
by: Huang, Wenyu, et al.
Published: (2025)
BookWorm: A Dataset for Character Description and Analysis
by: Papoudakis, Argyrios, et al.
Published: (2024)
by: Papoudakis, Argyrios, et al.
Published: (2024)
Think Before you Write: QA-Guided Reasoning for Character Descriptions in Books
by: Papoudakis, Argyrios, et al.
Published: (2026)
by: Papoudakis, Argyrios, et al.
Published: (2026)
Meta-Adaptive Prompt Distillation for Few-Shot Visual Question Answering
by: Gupta, Akash, et al.
Published: (2025)
by: Gupta, Akash, et al.
Published: (2025)
Explanatory Summarization with Discourse-Driven Planning
by: Liu, Dongqi, et al.
Published: (2025)
by: Liu, Dongqi, et al.
Published: (2025)
SimLM: Can Language Models Infer Parameters of Physical Systems?
by: Memery, Sean, et al.
Published: (2023)
by: Memery, Sean, et al.
Published: (2023)
ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning
by: Chen, Mingyang, et al.
Published: (2025)
by: Chen, Mingyang, et al.
Published: (2025)
Consistency-Aware Editing for Entity-level Unlearning in Language Models
by: Han, Xiaoqi, et al.
Published: (2025)
by: Han, Xiaoqi, et al.
Published: (2025)
Rethinking Memory in LLM based Agents: Representations, Operations, and Emerging Topics
by: Du, Yiming, et al.
Published: (2025)
by: Du, Yiming, et al.
Published: (2025)
Prompting Large Language Models with Knowledge Graphs for Question Answering Involving Long-tail Facts
by: Huang, Wenyu, et al.
Published: (2024)
by: Huang, Wenyu, et al.
Published: (2024)
Improving the Reliability of LLMs: Combining CoT, RAG, Self-Consistency, and Self-Verification
by: Kumar, Adarsh, et al.
Published: (2025)
by: Kumar, Adarsh, et al.
Published: (2025)
Scientific Knowledge-driven Decoding Constraints Improving the Reliability of LLMs
by: Ma, Maotian, et al.
Published: (2026)
by: Ma, Maotian, et al.
Published: (2026)
Evaluating and Safeguarding the Adversarial Robustness of Retrieval-Based In-Context Learning
by: Yu, Simon, et al.
Published: (2024)
by: Yu, Simon, et al.
Published: (2024)
How Reliable Are Automatic Evaluation Methods for Instruction-Tuned LLMs?
by: Doostmohammadi, Ehsan, et al.
Published: (2024)
by: Doostmohammadi, Ehsan, et al.
Published: (2024)
Low-Rank Adaptation for Multilingual Summarization: An Empirical Study
by: Whitehouse, Chenxi, et al.
Published: (2023)
by: Whitehouse, Chenxi, et al.
Published: (2023)
ScreenWriter: Automatic Screenplay Generation and Movie Summarisation
by: Mahon, Louis, et al.
Published: (2024)
by: Mahon, Louis, et al.
Published: (2024)
On Robustness and Reliability of Benchmark-Based Evaluation of LLMs
by: Lunardi, Riccardo, et al.
Published: (2025)
by: Lunardi, Riccardo, et al.
Published: (2025)
<think> So let's replace this phrase with insult... </think> Lessons learned from generation of toxic texts with LLMs
by: Pletenev, Sergey, et al.
Published: (2025)
by: Pletenev, Sergey, et al.
Published: (2025)
XplainLLM: A Knowledge-Augmented Dataset for Reliable Grounded Explanations in LLMs
by: Chen, Zichen, et al.
Published: (2023)
by: Chen, Zichen, et al.
Published: (2023)
Double-Calibration: Towards Reliable LLMs via Calibrating Knowledge and Reasoning Confidence
by: Lu, Yuyin, et al.
Published: (2026)
by: Lu, Yuyin, et al.
Published: (2026)
ReMA: Learning to Meta-think for LLMs with Multi-Agent Reinforcement Learning
by: Wan, Ziyu, et al.
Published: (2025)
by: Wan, Ziyu, et al.
Published: (2025)
Beneath the Surface of Consistency: Exploring Cross-lingual Knowledge Representation Sharing in LLMs
by: Ifergan, Maxim, et al.
Published: (2024)
by: Ifergan, Maxim, et al.
Published: (2024)
An Extensive Evaluation of PDDL Capabilities in off-the-shelf LLMs
by: Vyas, Kaustubh, et al.
Published: (2025)
by: Vyas, Kaustubh, et al.
Published: (2025)
Memorization and Knowledge Injection in Gated LLMs
by: Pan, Xu, et al.
Published: (2025)
by: Pan, Xu, et al.
Published: (2025)
LiveCLKTBench: Towards Reliable Evaluation of Cross-Lingual Knowledge Transfer in Multilingual LLMs
by: Guo, Pei-Fu, et al.
Published: (2025)
by: Guo, Pei-Fu, et al.
Published: (2025)
Noise-powered Multi-modal Knowledge Graph Representation Framework
by: Chen, Zhuo, et al.
Published: (2024)
by: Chen, Zhuo, et al.
Published: (2024)
How to Make the Most of LLMs' Grammatical Knowledge for Acceptability Judgments
by: Ide, Yusuke, et al.
Published: (2024)
by: Ide, Yusuke, et al.
Published: (2024)
What Is That Talk About? A Video-to-Text Summarization Dataset for Scientific Presentations
by: Liu, Dongqi, et al.
Published: (2025)
by: Liu, Dongqi, et al.
Published: (2025)
The Reliability of LLMs for Medical Diagnosis: An Examination of Consistency, Manipulation, and Contextual Awareness
by: Subedi, Krishna
Published: (2025)
by: Subedi, Krishna
Published: (2025)
COTET: Cross-view Optimal Transport for Knowledge Graph Entity Typing
by: Hu, Zhiwei, et al.
Published: (2024)
by: Hu, Zhiwei, et al.
Published: (2024)
Distilling Reasoning Without Knowledge: A Framework for Reliable LLMs
by: Kietkajornrit, Auksarapak, et al.
Published: (2026)
by: Kietkajornrit, Auksarapak, et al.
Published: (2026)
Similar Items
-
Long-Form Information Alignment Evaluation Beyond Atomic Facts
by: Zheng, Danna, et al.
Published: (2025) -
Archer: A Human-Labeled Text-to-SQL Dataset with Arithmetic, Commonsense and Hypothetical Reasoning
by: Zheng, Danna, et al.
Published: (2024) -
TrustScore: Reference-Free Evaluation of LLM Response Trustworthiness
by: Zheng, Danna, et al.
Published: (2024) -
Reasoning about Intent for Ambiguous Requests
by: Saparina, Irina, et al.
Published: (2025) -
Debating for Better Reasoning: An Unsupervised Multimodal Approach
by: Adhikari, Ashutosh, et al.
Published: (2025)