SimpleQA Verified: A Reliable Factuality Benchmark to Measure Parametric Knowledge
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Haas, Lukas, Yona, Gal, D'Antonio, Giovanni, Goldshtein, Sasha, Das, Dipanjan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Chinese SimpleQA: A Chinese Factuality Evaluation for Large Language Models
von: He, Yancheng, et al.
Veröffentlicht: (2024)
von: He, Yancheng, et al.
Veröffentlicht: (2024)
DeepSearchQA: Bridging the Comprehensiveness Gap for Deep Research Agents
von: Gupta, Nikita, et al.
Veröffentlicht: (2026)
von: Gupta, Nikita, et al.
Veröffentlicht: (2026)
Empty Shelves or Lost Keys? Recall Is the Bottleneck for Parametric Factuality
von: Calderon, Nitay, et al.
Veröffentlicht: (2026)
von: Calderon, Nitay, et al.
Veröffentlicht: (2026)
Video SimpleQA: Towards Factuality Evaluation in Large Video Language Models
von: Cao, Meng, et al.
Veröffentlicht: (2025)
von: Cao, Meng, et al.
Veröffentlicht: (2025)
KoSimpleQA: A Korean Factuality Benchmark with an Analysis of Reasoning LLMs
von: Ko, Donghyeon, et al.
Veröffentlicht: (2025)
von: Ko, Donghyeon, et al.
Veröffentlicht: (2025)
Narrowing the Knowledge Evaluation Gap: Open-Domain Question Answering with Multi-Granularity Answers
von: Yona, Gal, et al.
Veröffentlicht: (2024)
von: Yona, Gal, et al.
Veröffentlicht: (2024)
CodeSimpleQA: Scaling Factuality in Code Large Language Models
von: Yang, Jian, et al.
Veröffentlicht: (2025)
von: Yang, Jian, et al.
Veröffentlicht: (2025)
Factuality or Fiction? Benchmarking Modern LLMs on Ambiguous QA with Citations
von: Patel, Maya, et al.
Veröffentlicht: (2024)
von: Patel, Maya, et al.
Veröffentlicht: (2024)
AdversaRiskQA: An Adversarial Factuality Benchmark for High-Risk Domains
von: Szelestey, Adam, et al.
Veröffentlicht: (2026)
von: Szelestey, Adam, et al.
Veröffentlicht: (2026)
Can Large Language Models Faithfully Express Their Intrinsic Uncertainty in Words?
von: Yona, Gal, et al.
Veröffentlicht: (2024)
von: Yona, Gal, et al.
Veröffentlicht: (2024)
Hallucinations Undermine Trust; Metacognition is a Way Forward
von: Yona, Gal, et al.
Veröffentlicht: (2026)
von: Yona, Gal, et al.
Veröffentlicht: (2026)
ComparisonQA: Evaluating Factuality Robustness of LLMs Through Knowledge Frequency Control and Uncertainty
von: Zong, Qing, et al.
Veröffentlicht: (2024)
von: Zong, Qing, et al.
Veröffentlicht: (2024)
A Diagnostic Benchmark for Sweden-Related Factual Knowledge
von: Kunz, Jenny
Veröffentlicht: (2025)
von: Kunz, Jenny
Veröffentlicht: (2025)
DyKnow: Dynamically Verifying Time-Sensitive Factual Knowledge in LLMs
von: Mousavi, Seyed Mahed, et al.
Veröffentlicht: (2024)
von: Mousavi, Seyed Mahed, et al.
Veröffentlicht: (2024)
TACT: Advancing Complex Aggregative Reasoning with Information Extraction Tools
von: Caciularu, Avi, et al.
Veröffentlicht: (2024)
von: Caciularu, Avi, et al.
Veröffentlicht: (2024)
The FACTS Leaderboard: A Comprehensive Benchmark for Large Language Model Factuality
von: Cheng, Aileen, et al.
Veröffentlicht: (2025)
von: Cheng, Aileen, et al.
Veröffentlicht: (2025)
Chinese SafetyQA: A Safety Short-form Factuality Benchmark for Large Language Models
von: Tan, Yingshui, et al.
Veröffentlicht: (2024)
von: Tan, Yingshui, et al.
Veröffentlicht: (2024)
Stop When Reasoning Converges: Semantic-Preserving Early Exit for Reasoning Models
von: Min, Dehai, et al.
Veröffentlicht: (2026)
von: Min, Dehai, et al.
Veröffentlicht: (2026)
Towards Reliable Latent Knowledge Estimation in LLMs: Zero-Prompt Many-Shot Based Factual Knowledge Extraction
von: Wu, Qinyuan, et al.
Veröffentlicht: (2024)
von: Wu, Qinyuan, et al.
Veröffentlicht: (2024)
Keep Guessing? When Considering Inference Scaling, Mind the Baselines
von: Yona, Gal, et al.
Veröffentlicht: (2024)
von: Yona, Gal, et al.
Veröffentlicht: (2024)
Scientific QA System with Verifiable Answers
von: Ljajić, Adela, et al.
Veröffentlicht: (2024)
von: Ljajić, Adela, et al.
Veröffentlicht: (2024)
Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?
von: Gekhman, Zorik, et al.
Veröffentlicht: (2024)
von: Gekhman, Zorik, et al.
Veröffentlicht: (2024)
MoleculeQA: A Dataset to Evaluate Factual Accuracy in Molecular Comprehension
von: Lu, Xingyu, et al.
Veröffentlicht: (2024)
von: Lu, Xingyu, et al.
Veröffentlicht: (2024)
Understanding QA generation: Extracting Parametric and Contextual Knowledge with CQA for Low Resource Bangla Language
von: Azmary, Umme Abira, et al.
Veröffentlicht: (2026)
von: Azmary, Umme Abira, et al.
Veröffentlicht: (2026)
The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input
von: Jacovi, Alon, et al.
Veröffentlicht: (2025)
von: Jacovi, Alon, et al.
Veröffentlicht: (2025)
Factual Knowledge in Language Models: Robustness and Anomalies under Simple Temporal Context Variations
von: Khodja, Hichem Ammar, et al.
Veröffentlicht: (2025)
von: Khodja, Hichem Ammar, et al.
Veröffentlicht: (2025)
Verify with Caution: The Pitfalls of Relying on Imperfect Factuality Metrics
von: Godbole, Ameya, et al.
Veröffentlicht: (2025)
von: Godbole, Ameya, et al.
Veröffentlicht: (2025)
BioPulse-QA: A Dynamic Biomedical Question-Answering Benchmark for Evaluating Factuality, Robustness, and Bias in Large Language Models
von: Bhattarai, Kriti, et al.
Veröffentlicht: (2026)
von: Bhattarai, Kriti, et al.
Veröffentlicht: (2026)
DiVA: Fine-grained Factuality Verification with Agentic-Discriminative Verifier
von: Huang, Hui, et al.
Veröffentlicht: (2026)
von: Huang, Hui, et al.
Veröffentlicht: (2026)
Do Automatic Factuality Metrics Measure Factuality? A Critical Evaluation
von: Ramprasad, Sanjana, et al.
Veröffentlicht: (2024)
von: Ramprasad, Sanjana, et al.
Veröffentlicht: (2024)
Iterate Until Retrieved: Factual Nugget Optimization for Discoverable Continual Corrections in Agentic RAG
von: Hazoom, Moshe, et al.
Veröffentlicht: (2026)
von: Hazoom, Moshe, et al.
Veröffentlicht: (2026)
Towards Verifiable Generation: A Benchmark for Knowledge-aware Language Model Attribution
von: Li, Xinze, et al.
Veröffentlicht: (2023)
von: Li, Xinze, et al.
Veröffentlicht: (2023)
SimpleVQA: Multimodal Factuality Evaluation for Multimodal Large Language Models
von: Cheng, Xianfu, et al.
Veröffentlicht: (2025)
von: Cheng, Xianfu, et al.
Veröffentlicht: (2025)
FACTORY: A Challenging Human-Verified Prompt Set for Long-Form Factuality
von: Chen, Mingda, et al.
Veröffentlicht: (2025)
von: Chen, Mingda, et al.
Veröffentlicht: (2025)
Can LLMs Evaluate Complex Attribution in QA? Automatic Benchmarking using Knowledge Graphs
von: Hu, Nan, et al.
Veröffentlicht: (2024)
von: Hu, Nan, et al.
Veröffentlicht: (2024)
Inside-Out: Hidden Factual Knowledge in LLMs
von: Gekhman, Zorik, et al.
Veröffentlicht: (2025)
von: Gekhman, Zorik, et al.
Veröffentlicht: (2025)
LLMs as Repositories of Factual Knowledge: Limitations and Solutions
von: Mousavi, Seyed Mahed, et al.
Veröffentlicht: (2025)
von: Mousavi, Seyed Mahed, et al.
Veröffentlicht: (2025)
Tracing Multilingual Factual Knowledge Acquisition in Pretraining
von: Liu, Yihong, et al.
Veröffentlicht: (2025)
von: Liu, Yihong, et al.
Veröffentlicht: (2025)
UniKnow: A Unified Framework for Reliable Language Model Behavior across Parametric and External Knowledge
von: Kim, Youna, et al.
Veröffentlicht: (2025)
von: Kim, Youna, et al.
Veröffentlicht: (2025)
BioGraphletQA: Knowledge-Anchored Generation of Complex QA Datasets
von: Jonker, Richard A. A., et al.
Veröffentlicht: (2026)
von: Jonker, Richard A. A., et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Chinese SimpleQA: A Chinese Factuality Evaluation for Large Language Models
von: He, Yancheng, et al.
Veröffentlicht: (2024) -
DeepSearchQA: Bridging the Comprehensiveness Gap for Deep Research Agents
von: Gupta, Nikita, et al.
Veröffentlicht: (2026) -
Empty Shelves or Lost Keys? Recall Is the Bottleneck for Parametric Factuality
von: Calderon, Nitay, et al.
Veröffentlicht: (2026) -
Video SimpleQA: Towards Factuality Evaluation in Large Video Language Models
von: Cao, Meng, et al.
Veröffentlicht: (2025) -
KoSimpleQA: A Korean Factuality Benchmark with an Analysis of Reasoning LLMs
von: Ko, Donghyeon, et al.
Veröffentlicht: (2025)