SimpleQA Verified: A Reliable Factuality Benchmark to Measure Parametric Knowledge
Fuente:
arXiv
Saved in:
| Main Authors: | Haas, Lukas, Yona, Gal, D'Antonio, Giovanni, Goldshtein, Sasha, Das, Dipanjan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Chinese SimpleQA: A Chinese Factuality Evaluation for Large Language Models
by: He, Yancheng, et al.
Published: (2024)
by: He, Yancheng, et al.
Published: (2024)
DeepSearchQA: Bridging the Comprehensiveness Gap for Deep Research Agents
by: Gupta, Nikita, et al.
Published: (2026)
by: Gupta, Nikita, et al.
Published: (2026)
Empty Shelves or Lost Keys? Recall Is the Bottleneck for Parametric Factuality
by: Calderon, Nitay, et al.
Published: (2026)
by: Calderon, Nitay, et al.
Published: (2026)
Video SimpleQA: Towards Factuality Evaluation in Large Video Language Models
by: Cao, Meng, et al.
Published: (2025)
by: Cao, Meng, et al.
Published: (2025)
KoSimpleQA: A Korean Factuality Benchmark with an Analysis of Reasoning LLMs
by: Ko, Donghyeon, et al.
Published: (2025)
by: Ko, Donghyeon, et al.
Published: (2025)
Narrowing the Knowledge Evaluation Gap: Open-Domain Question Answering with Multi-Granularity Answers
by: Yona, Gal, et al.
Published: (2024)
by: Yona, Gal, et al.
Published: (2024)
CodeSimpleQA: Scaling Factuality in Code Large Language Models
by: Yang, Jian, et al.
Published: (2025)
by: Yang, Jian, et al.
Published: (2025)
Factuality or Fiction? Benchmarking Modern LLMs on Ambiguous QA with Citations
by: Patel, Maya, et al.
Published: (2024)
by: Patel, Maya, et al.
Published: (2024)
AdversaRiskQA: An Adversarial Factuality Benchmark for High-Risk Domains
by: Szelestey, Adam, et al.
Published: (2026)
by: Szelestey, Adam, et al.
Published: (2026)
Can Large Language Models Faithfully Express Their Intrinsic Uncertainty in Words?
by: Yona, Gal, et al.
Published: (2024)
by: Yona, Gal, et al.
Published: (2024)
Hallucinations Undermine Trust; Metacognition is a Way Forward
by: Yona, Gal, et al.
Published: (2026)
by: Yona, Gal, et al.
Published: (2026)
ComparisonQA: Evaluating Factuality Robustness of LLMs Through Knowledge Frequency Control and Uncertainty
by: Zong, Qing, et al.
Published: (2024)
by: Zong, Qing, et al.
Published: (2024)
A Diagnostic Benchmark for Sweden-Related Factual Knowledge
by: Kunz, Jenny
Published: (2025)
by: Kunz, Jenny
Published: (2025)
DyKnow: Dynamically Verifying Time-Sensitive Factual Knowledge in LLMs
by: Mousavi, Seyed Mahed, et al.
Published: (2024)
by: Mousavi, Seyed Mahed, et al.
Published: (2024)
TACT: Advancing Complex Aggregative Reasoning with Information Extraction Tools
by: Caciularu, Avi, et al.
Published: (2024)
by: Caciularu, Avi, et al.
Published: (2024)
The FACTS Leaderboard: A Comprehensive Benchmark for Large Language Model Factuality
by: Cheng, Aileen, et al.
Published: (2025)
by: Cheng, Aileen, et al.
Published: (2025)
Chinese SafetyQA: A Safety Short-form Factuality Benchmark for Large Language Models
by: Tan, Yingshui, et al.
Published: (2024)
by: Tan, Yingshui, et al.
Published: (2024)
Stop When Reasoning Converges: Semantic-Preserving Early Exit for Reasoning Models
by: Min, Dehai, et al.
Published: (2026)
by: Min, Dehai, et al.
Published: (2026)
Towards Reliable Latent Knowledge Estimation in LLMs: Zero-Prompt Many-Shot Based Factual Knowledge Extraction
by: Wu, Qinyuan, et al.
Published: (2024)
by: Wu, Qinyuan, et al.
Published: (2024)
Keep Guessing? When Considering Inference Scaling, Mind the Baselines
by: Yona, Gal, et al.
Published: (2024)
by: Yona, Gal, et al.
Published: (2024)
Scientific QA System with Verifiable Answers
by: Ljajić, Adela, et al.
Published: (2024)
by: Ljajić, Adela, et al.
Published: (2024)
Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?
by: Gekhman, Zorik, et al.
Published: (2024)
by: Gekhman, Zorik, et al.
Published: (2024)
MoleculeQA: A Dataset to Evaluate Factual Accuracy in Molecular Comprehension
by: Lu, Xingyu, et al.
Published: (2024)
by: Lu, Xingyu, et al.
Published: (2024)
Understanding QA generation: Extracting Parametric and Contextual Knowledge with CQA for Low Resource Bangla Language
by: Azmary, Umme Abira, et al.
Published: (2026)
by: Azmary, Umme Abira, et al.
Published: (2026)
The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input
by: Jacovi, Alon, et al.
Published: (2025)
by: Jacovi, Alon, et al.
Published: (2025)
Factual Knowledge in Language Models: Robustness and Anomalies under Simple Temporal Context Variations
by: Khodja, Hichem Ammar, et al.
Published: (2025)
by: Khodja, Hichem Ammar, et al.
Published: (2025)
Verify with Caution: The Pitfalls of Relying on Imperfect Factuality Metrics
by: Godbole, Ameya, et al.
Published: (2025)
by: Godbole, Ameya, et al.
Published: (2025)
BioPulse-QA: A Dynamic Biomedical Question-Answering Benchmark for Evaluating Factuality, Robustness, and Bias in Large Language Models
by: Bhattarai, Kriti, et al.
Published: (2026)
by: Bhattarai, Kriti, et al.
Published: (2026)
DiVA: Fine-grained Factuality Verification with Agentic-Discriminative Verifier
by: Huang, Hui, et al.
Published: (2026)
by: Huang, Hui, et al.
Published: (2026)
Do Automatic Factuality Metrics Measure Factuality? A Critical Evaluation
by: Ramprasad, Sanjana, et al.
Published: (2024)
by: Ramprasad, Sanjana, et al.
Published: (2024)
Iterate Until Retrieved: Factual Nugget Optimization for Discoverable Continual Corrections in Agentic RAG
by: Hazoom, Moshe, et al.
Published: (2026)
by: Hazoom, Moshe, et al.
Published: (2026)
Towards Verifiable Generation: A Benchmark for Knowledge-aware Language Model Attribution
by: Li, Xinze, et al.
Published: (2023)
by: Li, Xinze, et al.
Published: (2023)
SimpleVQA: Multimodal Factuality Evaluation for Multimodal Large Language Models
by: Cheng, Xianfu, et al.
Published: (2025)
by: Cheng, Xianfu, et al.
Published: (2025)
FACTORY: A Challenging Human-Verified Prompt Set for Long-Form Factuality
by: Chen, Mingda, et al.
Published: (2025)
by: Chen, Mingda, et al.
Published: (2025)
Can LLMs Evaluate Complex Attribution in QA? Automatic Benchmarking using Knowledge Graphs
by: Hu, Nan, et al.
Published: (2024)
by: Hu, Nan, et al.
Published: (2024)
Inside-Out: Hidden Factual Knowledge in LLMs
by: Gekhman, Zorik, et al.
Published: (2025)
by: Gekhman, Zorik, et al.
Published: (2025)
LLMs as Repositories of Factual Knowledge: Limitations and Solutions
by: Mousavi, Seyed Mahed, et al.
Published: (2025)
by: Mousavi, Seyed Mahed, et al.
Published: (2025)
Tracing Multilingual Factual Knowledge Acquisition in Pretraining
by: Liu, Yihong, et al.
Published: (2025)
by: Liu, Yihong, et al.
Published: (2025)
UniKnow: A Unified Framework for Reliable Language Model Behavior across Parametric and External Knowledge
by: Kim, Youna, et al.
Published: (2025)
by: Kim, Youna, et al.
Published: (2025)
BioGraphletQA: Knowledge-Anchored Generation of Complex QA Datasets
by: Jonker, Richard A. A., et al.
Published: (2026)
by: Jonker, Richard A. A., et al.
Published: (2026)
Similar Items
-
Chinese SimpleQA: A Chinese Factuality Evaluation for Large Language Models
by: He, Yancheng, et al.
Published: (2024) -
DeepSearchQA: Bridging the Comprehensiveness Gap for Deep Research Agents
by: Gupta, Nikita, et al.
Published: (2026) -
Empty Shelves or Lost Keys? Recall Is the Bottleneck for Parametric Factuality
by: Calderon, Nitay, et al.
Published: (2026) -
Video SimpleQA: Towards Factuality Evaluation in Large Video Language Models
by: Cao, Meng, et al.
Published: (2025) -
KoSimpleQA: A Korean Factuality Benchmark with an Analysis of Reasoning LLMs
by: Ko, Donghyeon, et al.
Published: (2025)