ComparisonQA: Evaluating Factuality Robustness of LLMs Through Knowledge Frequency Control and Uncertainty
Fuente:
arXiv
Saved in:
| Main Authors: | Zong, Qing, Wang, Zhaowei, Zheng, Tianshi, Ren, Xiyu, Song, Yangqiu |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
KnowShiftQA: How Robust are RAG Systems when Textbook Knowledge Shifts in K-12 Education?
by: Zheng, Tianshi, et al.
Published: (2024)
by: Zheng, Tianshi, et al.
Published: (2024)
KNOWCOMP POKEMON Team at DialAM-2024: A Two-Stage Pipeline for Detecting Relations in Dialogical Argument Mining
by: Zheng, Zihao, et al.
Published: (2024)
by: Zheng, Zihao, et al.
Published: (2024)
CLR-Fact: Evaluating the Complex Logical Reasoning Capability of Large Language Models over Factual Knowledge
by: Zheng, Tianshi, et al.
Published: (2024)
by: Zheng, Tianshi, et al.
Published: (2024)
CritiCal: Can Critique Help LLM Uncertainty or Confidence Calibration?
by: Zong, Qing, et al.
Published: (2025)
by: Zong, Qing, et al.
Published: (2025)
Revisiting Epistemic Markers in Confidence Estimation: Can Markers Accurately Reflect Large Language Models' Uncertainty?
by: Liu, Jiayu, et al.
Published: (2025)
by: Liu, Jiayu, et al.
Published: (2025)
NAACL: Noise-AwAre Verbal Confidence Calibration for Robust LLMs in RAG Systems
by: Liu, Jiayu, et al.
Published: (2026)
by: Liu, Jiayu, et al.
Published: (2026)
Concept-Reversed Winograd Schema Challenge: Evaluating and Improving Robust Reasoning in Large Language Models via Abstraction
by: Han, Kaiqiao, et al.
Published: (2024)
by: Han, Kaiqiao, et al.
Published: (2024)
INFERENCEDYNAMICS: Efficient Routing Across LLMs through Structured Capability and Knowledge Profiling
by: Shi, Haochen, et al.
Published: (2025)
by: Shi, Haochen, et al.
Published: (2025)
The Cognitive Bandwidth Bottleneck: Shifting Long-Horizon Agent from Planning with Actions to Planning with Schemas
by: Xu, Baixuan, et al.
Published: (2025)
by: Xu, Baixuan, et al.
Published: (2025)
Advancing Abductive Reasoning in Knowledge Graphs through Complex Logical Hypothesis Generation
by: Bai, Jiaxin, et al.
Published: (2023)
by: Bai, Jiaxin, et al.
Published: (2023)
CodeGraph: Enhancing Graph Reasoning of LLMs with Code
by: Cai, Qiaolong, et al.
Published: (2024)
by: Cai, Qiaolong, et al.
Published: (2024)
Structuring the Unstructured: A Systematic Review of Text-to-Structure Generation for Agentic AI with a Universal Evaluation Framework
by: Deng, Zheye, et al.
Published: (2025)
by: Deng, Zheye, et al.
Published: (2025)
LLM-Hanabi: Evaluating Multi-Agent Gameplays with Theory-of-Mind and Rationale Inference in Imperfect Information Collaboration Game
by: Liang, Fangzhou, et al.
Published: (2025)
by: Liang, Fangzhou, et al.
Published: (2025)
Evaluating and Enhancing LLMs Agent based on Theory of Mind in Guandan: A Multi-Player Cooperative Game under Imperfect Information
by: Yim, Yauwai, et al.
Published: (2024)
by: Yim, Yauwai, et al.
Published: (2024)
AbsInstruct: Eliciting Abstraction Ability from LLMs through Explanation Tuning with Plausibility Estimation
by: Wang, Zhaowei, et al.
Published: (2024)
by: Wang, Zhaowei, et al.
Published: (2024)
CKBP v2: Better Annotation and Reasoning for Commonsense Knowledge Base Population
by: Fang, Tianqing, et al.
Published: (2023)
by: Fang, Tianqing, et al.
Published: (2023)
Factuality or Fiction? Benchmarking Modern LLMs on Ambiguous QA with Citations
by: Patel, Maya, et al.
Published: (2024)
by: Patel, Maya, et al.
Published: (2024)
Chinese SimpleQA: A Chinese Factuality Evaluation for Large Language Models
by: He, Yancheng, et al.
Published: (2024)
by: He, Yancheng, et al.
Published: (2024)
ConstraintChecker: A Plugin for Large Language Models to Reason on Commonsense Knowledge Bases
by: Do, Quyet V., et al.
Published: (2024)
by: Do, Quyet V., et al.
Published: (2024)
LLMs are Frequency Pattern Learners in Natural Language Inference
by: Cheng, Liang, et al.
Published: (2025)
by: Cheng, Liang, et al.
Published: (2025)
MoleculeQA: A Dataset to Evaluate Factual Accuracy in Molecular Comprehension
by: Lu, Xingyu, et al.
Published: (2024)
by: Lu, Xingyu, et al.
Published: (2024)
UAQFact: Evaluating Factual Knowledge Utilization of LLMs on Unanswerable Questions
by: Tan, Chuanyuan, et al.
Published: (2025)
by: Tan, Chuanyuan, et al.
Published: (2025)
What Really is Commonsense Knowledge?
by: Do, Quyet V., et al.
Published: (2024)
by: Do, Quyet V., et al.
Published: (2024)
Getting Sick After Seeing a Doctor? Diagnosing and Mitigating Knowledge Conflicts in Event Temporal Reasoning
by: Fang, Tianqing, et al.
Published: (2023)
by: Fang, Tianqing, et al.
Published: (2023)
KoSimpleQA: A Korean Factuality Benchmark with an Analysis of Reasoning LLMs
by: Ko, Donghyeon, et al.
Published: (2025)
by: Ko, Donghyeon, et al.
Published: (2025)
SimpleQA Verified: A Reliable Factuality Benchmark to Measure Parametric Knowledge
by: Haas, Lukas, et al.
Published: (2025)
by: Haas, Lukas, et al.
Published: (2025)
The Curse of CoT: On the Limitations of Chain-of-Thought in In-Context Learning
by: Zheng, Tianshi, et al.
Published: (2025)
by: Zheng, Tianshi, et al.
Published: (2025)
MARS: Benchmarking the Metaphysical Reasoning Abilities of Language Models with a Multi-task Evaluation Dataset
by: Wang, Weiqi, et al.
Published: (2024)
by: Wang, Weiqi, et al.
Published: (2024)
BioPulse-QA: A Dynamic Biomedical Question-Answering Benchmark for Evaluating Factuality, Robustness, and Bias in Large Language Models
by: Bhattarai, Kriti, et al.
Published: (2026)
by: Bhattarai, Kriti, et al.
Published: (2026)
Measuring Aleatoric and Epistemic Uncertainty in LLMs: Empirical Evaluation on ID and OOD QA Tasks
by: Wang, Kevin, et al.
Published: (2025)
by: Wang, Kevin, et al.
Published: (2025)
Inside-Out: Hidden Factual Knowledge in LLMs
by: Gekhman, Zorik, et al.
Published: (2025)
by: Gekhman, Zorik, et al.
Published: (2025)
LLMs as Repositories of Factual Knowledge: Limitations and Solutions
by: Mousavi, Seyed Mahed, et al.
Published: (2025)
by: Mousavi, Seyed Mahed, et al.
Published: (2025)
Towards a Holistic Evaluation of LLMs on Factual Knowledge Recall
by: Yuan, Jiaqing, et al.
Published: (2024)
by: Yuan, Jiaqing, et al.
Published: (2024)
From Automation to Autonomy: A Survey on Large Language Models in Scientific Discovery
by: Zheng, Tianshi, et al.
Published: (2025)
by: Zheng, Tianshi, et al.
Published: (2025)
AutoGraph-R1: End-to-End Reinforcement Learning for Knowledge Graph Construction
by: Tsang, Hong Ting, et al.
Published: (2025)
by: Tsang, Hong Ting, et al.
Published: (2025)
Can LLMs Evaluate Complex Attribution in QA? Automatic Benchmarking using Knowledge Graphs
by: Hu, Nan, et al.
Published: (2024)
by: Hu, Nan, et al.
Published: (2024)
Think Through Uncertainty: Improving Long-Form Generation Factuality via Reasoning Calibration
by: Liu, Xin, et al.
Published: (2026)
by: Liu, Xin, et al.
Published: (2026)
Persuasion Tokens for Editing Factual Knowledge in LLMs
by: Youssef, Paul, et al.
Published: (2026)
by: Youssef, Paul, et al.
Published: (2026)
CodeSimpleQA: Scaling Factuality in Code Large Language Models
by: Yang, Jian, et al.
Published: (2025)
by: Yang, Jian, et al.
Published: (2025)
Causal Path Alignment: Anchoring the Optimization Trajectory for Controllable In-Parameter Knowledge Editing
by: Liu, Xiyu, et al.
Published: (2025)
by: Liu, Xiyu, et al.
Published: (2025)
Similar Items
-
KnowShiftQA: How Robust are RAG Systems when Textbook Knowledge Shifts in K-12 Education?
by: Zheng, Tianshi, et al.
Published: (2024) -
KNOWCOMP POKEMON Team at DialAM-2024: A Two-Stage Pipeline for Detecting Relations in Dialogical Argument Mining
by: Zheng, Zihao, et al.
Published: (2024) -
CLR-Fact: Evaluating the Complex Logical Reasoning Capability of Large Language Models over Factual Knowledge
by: Zheng, Tianshi, et al.
Published: (2024) -
CritiCal: Can Critique Help LLM Uncertainty or Confidence Calibration?
by: Zong, Qing, et al.
Published: (2025) -
Revisiting Epistemic Markers in Confidence Estimation: Can Markers Accurately Reflect Large Language Models' Uncertainty?
by: Liu, Jiayu, et al.
Published: (2025)