Towards a Holistic Evaluation of LLMs on Factual Knowledge Recall
Fuente:
arXiv
Saved in:
| Main Authors: | Yuan, Jiaqing, Pan, Lin, Hang, Chung-Wei, Guo, Jiang, Jiang, Jiarong, Min, Bonan, Ng, Patrick, Wang, Zhiguo |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PRACTIQ: A Practical Conversational Text-to-SQL dataset with Ambiguous and Unanswerable Queries
by: Dong, Mingwen, et al.
Published: (2024)
by: Dong, Mingwen, et al.
Published: (2024)
Unveiling Factual Recall Behaviors of Large Language Models through Knowledge Neurons
by: Wang, Yifei, et al.
Published: (2024)
by: Wang, Yifei, et al.
Published: (2024)
Locate-then-edit for Multi-hop Factual Recall under Knowledge Editing
by: Zhang, Zhuoran, et al.
Published: (2024)
by: Zhang, Zhuoran, et al.
Published: (2024)
Propagation and Pitfalls: Reasoning-based Assessment of Knowledge Editing through Counterfactual Tasks
by: Hua, Wenyue, et al.
Published: (2024)
by: Hua, Wenyue, et al.
Published: (2024)
Injecting Falsehoods: Adversarial Man-in-the-Middle Attacks Undermining Factual Recall in LLMs
by: Fastowski, Alina, et al.
Published: (2025)
by: Fastowski, Alina, et al.
Published: (2025)
Predictable Confabulations: Factual Recall by LLMs Scales with Model Size and Topic Frequency
by: Smith, Matthew L., et al.
Published: (2026)
by: Smith, Matthew L., et al.
Published: (2026)
Empty Shelves or Lost Keys? Recall Is the Bottleneck for Parametric Factuality
by: Calderon, Nitay, et al.
Published: (2026)
by: Calderon, Nitay, et al.
Published: (2026)
StratMem-Bench: Evaluating Strategic Memory Use in Virtual Character Conversation Beyond Factual Recall
by: Wu, Yerong, et al.
Published: (2026)
by: Wu, Yerong, et al.
Published: (2026)
Layerwise Recall and the Geometry of Interwoven Knowledge in LLMs
by: Lei, Ge, et al.
Published: (2025)
by: Lei, Ge, et al.
Published: (2025)
Towards a Holistic and Automated Evaluation Framework for Multi-Level Comprehension of LLMs in Book-Length Contexts
by: Deng, Jiaqi, et al.
Published: (2025)
by: Deng, Jiaqi, et al.
Published: (2025)
KG-FPQ: Evaluating Factuality Hallucination in LLMs with Knowledge Graph-based False Premise Questions
by: Zhu, Yanxu, et al.
Published: (2024)
by: Zhu, Yanxu, et al.
Published: (2024)
Can LLMs Refuse Questions They Do Not Know? Measuring Knowledge-Aware Refusal in Factual Tasks
by: Pan, Wenbo, et al.
Published: (2025)
by: Pan, Wenbo, et al.
Published: (2025)
WildHallucinations: Evaluating Long-form Factuality in LLMs with Real-World Entity Queries
by: Zhao, Wenting, et al.
Published: (2024)
by: Zhao, Wenting, et al.
Published: (2024)
Improving Factuality in LLMs via Inference-Time Knowledge Graph Construction
by: Wu, Shanglin, et al.
Published: (2025)
by: Wu, Shanglin, et al.
Published: (2025)
DyKnow: Dynamically Verifying Time-Sensitive Factual Knowledge in LLMs
by: Mousavi, Seyed Mahed, et al.
Published: (2024)
by: Mousavi, Seyed Mahed, et al.
Published: (2024)
Interpreting Key Mechanisms of Factual Recall in Transformer-Based Language Models
by: Lv, Ang, et al.
Published: (2024)
by: Lv, Ang, et al.
Published: (2024)
LiveCLKTBench: Towards Reliable Evaluation of Cross-Lingual Knowledge Transfer in Multilingual LLMs
by: Guo, Pei-Fu, et al.
Published: (2025)
by: Guo, Pei-Fu, et al.
Published: (2025)
MAD-Fact: A Multi-Agent Debate Framework for Long-Form Factuality Evaluation in LLMs
by: Ning, Yucheng, et al.
Published: (2025)
by: Ning, Yucheng, et al.
Published: (2025)
Face4RAG: Factual Consistency Evaluation for Retrieval Augmented Generation in Chinese
by: Xu, Yunqi, et al.
Published: (2024)
by: Xu, Yunqi, et al.
Published: (2024)
Can VLMs Recall Factual Associations From Visual References?
by: Ashok, Dhananjay, et al.
Published: (2025)
by: Ashok, Dhananjay, et al.
Published: (2025)
LLM Factoscope: Uncovering LLMs' Factual Discernment through Inner States Analysis
by: He, Jinwen, et al.
Published: (2023)
by: He, Jinwen, et al.
Published: (2023)
Self-Alignment for Factuality: Mitigating Hallucinations in LLMs via Self-Evaluation
by: Zhang, Xiaoying, et al.
Published: (2024)
by: Zhang, Xiaoying, et al.
Published: (2024)
MedReason: Eliciting Factual Medical Reasoning Steps in LLMs via Knowledge Graphs
by: Wu, Juncheng, et al.
Published: (2025)
by: Wu, Juncheng, et al.
Published: (2025)
FAME: Towards Factual Multi-Task Model Editing
by: Zeng, Li, et al.
Published: (2024)
by: Zeng, Li, et al.
Published: (2024)
Recall, Retrieve and Reason: Towards Better In-Context Relation Extraction
by: Li, Guozheng, et al.
Published: (2024)
by: Li, Guozheng, et al.
Published: (2024)
Towards Understanding Continual Factual Knowledge Acquisition of Language Models: From Theory to Algorithm
by: Wang, Haoyu, et al.
Published: (2026)
by: Wang, Haoyu, et al.
Published: (2026)
Do Automatic Factuality Metrics Measure Factuality? A Critical Evaluation
by: Ramprasad, Sanjana, et al.
Published: (2024)
by: Ramprasad, Sanjana, et al.
Published: (2024)
FoundaBench: Evaluating Chinese Fundamental Knowledge Capabilities of Large Language Models
by: Li, Wei, et al.
Published: (2024)
by: Li, Wei, et al.
Published: (2024)
WaterBench: Towards Holistic Evaluation of Watermarks for Large Language Models
by: Tu, Shangqing, et al.
Published: (2023)
by: Tu, Shangqing, et al.
Published: (2023)
Reasoning Models Hallucinate More: Factuality-Aware Reinforcement Learning for Large Reasoning Models
by: Li, Junyi, et al.
Published: (2025)
by: Li, Junyi, et al.
Published: (2025)
Is Factuality Enhancement a Free Lunch For LLMs? Better Factuality Can Lead to Worse Context-Faithfulness
by: Bi, Baolong, et al.
Published: (2024)
by: Bi, Baolong, et al.
Published: (2024)
Knowledgeable In-Context Tuning: Exploring and Exploiting Factual Knowledge for In-Context Learning
by: Wang, Jianing, et al.
Published: (2023)
by: Wang, Jianing, et al.
Published: (2023)
Editing Factual Knowledge and Explanatory Ability of Medical Large Language Models
by: Xu, Derong, et al.
Published: (2024)
by: Xu, Derong, et al.
Published: (2024)
Aligning Knowledge Graphs and Language Models for Factual Accuracy
by: Nishat, Nur A Zarin, et al.
Published: (2025)
by: Nishat, Nur A Zarin, et al.
Published: (2025)
MM-CRITIC: A Holistic Evaluation of Large Multimodal Models as Multimodal Critique
by: Zeng, Gailun, et al.
Published: (2025)
by: Zeng, Gailun, et al.
Published: (2025)
Selective Expert Guidance for Effective and Diverse Exploration in Reinforcement Learning of LLMs
by: Jiang, Zishang, et al.
Published: (2025)
by: Jiang, Zishang, et al.
Published: (2025)
CultureScope: A Dimensional Lens for Probing Cultural Understanding in LLMs
by: Zhang, Jinghao, et al.
Published: (2025)
by: Zhang, Jinghao, et al.
Published: (2025)
CiteEval: Principle-Driven Citation Evaluation for Source Attribution
by: Xu, Yumo, et al.
Published: (2025)
by: Xu, Yumo, et al.
Published: (2025)
Generating Benchmarks for Factuality Evaluation of Language Models
by: Muhlgay, Dor, et al.
Published: (2023)
by: Muhlgay, Dor, et al.
Published: (2023)
Evaluating the Factuality of Large Language Models using Large-Scale Knowledge Graphs
by: Liu, Xiaoze, et al.
Published: (2024)
by: Liu, Xiaoze, et al.
Published: (2024)
Similar Items
-
PRACTIQ: A Practical Conversational Text-to-SQL dataset with Ambiguous and Unanswerable Queries
by: Dong, Mingwen, et al.
Published: (2024) -
Unveiling Factual Recall Behaviors of Large Language Models through Knowledge Neurons
by: Wang, Yifei, et al.
Published: (2024) -
Locate-then-edit for Multi-hop Factual Recall under Knowledge Editing
by: Zhang, Zhuoran, et al.
Published: (2024) -
Propagation and Pitfalls: Reasoning-based Assessment of Knowledge Editing through Counterfactual Tasks
by: Hua, Wenyue, et al.
Published: (2024) -
Injecting Falsehoods: Adversarial Man-in-the-Middle Attacks Undermining Factual Recall in LLMs
by: Fastowski, Alina, et al.
Published: (2025)