OpenFactCheck: A Unified Framework for Factuality Evaluation of LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Iqbal, Hasan, Wang, Yuxia, Wang, Minghan, Georgiev, Georgi, Geng, Jiahui, Gurevych, Iryna, Nakov, Preslav |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
OpenFactCheck: Building, Benchmarking Customized Fact-Checking Systems and Evaluating the Factuality of Claims and LLMs
by: Wang, Yuxia, et al.
Published: (2024)
by: Wang, Yuxia, et al.
Published: (2024)
UrduFactCheck: An Agentic Fact-Checking Framework for Urdu with Evidence Boosting and Benchmarking
by: Ahmad, Sarfraz, et al.
Published: (2025)
by: Ahmad, Sarfraz, et al.
Published: (2025)
The CLEF-2025 CheckThat! Lab: Subjectivity, Fact-Checking, Claim Normalization, and Retrieval
by: Alam, Firoj, et al.
Published: (2025)
by: Alam, Firoj, et al.
Published: (2025)
Retrieve-Refine-Calibrate: A Framework for Complex Claim Fact-Checking
by: Sun, Mingwei, et al.
Published: (2026)
by: Sun, Mingwei, et al.
Published: (2026)
Integrating Causal Reasoning into Automated Fact-Checking
by: Rebboud, Youssra, et al.
Published: (2025)
by: Rebboud, Youssra, et al.
Published: (2025)
Multimodal Large Language Models to Support Real-World Fact-Checking
by: Geng, Jiahui, et al.
Published: (2024)
by: Geng, Jiahui, et al.
Published: (2024)
A Knowledge Enhanced Learning and Semantic Composition Model for Multi-Claim Fact Checking
by: Wang, Shuai, et al.
Published: (2021)
by: Wang, Shuai, et al.
Published: (2021)
Comprehensiveness Metrics for Automatic Evaluation of Factual Recall in Text Generation
by: Dejl, Adam, et al.
Published: (2025)
by: Dejl, Adam, et al.
Published: (2025)
GenAI Content Detection Task 3: Cross-Domain Machine-Generated Text Detection Challenge
by: Dugan, Liam, et al.
Published: (2025)
by: Dugan, Liam, et al.
Published: (2025)
Mr. Snuffleupagus at SemEval-2025 Task 4: Unlearning Factual Knowledge from LLMs Using Adaptive RMU
by: Dosajh, Arjun, et al.
Published: (2025)
by: Dosajh, Arjun, et al.
Published: (2025)
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
by: Saji, Alan, et al.
Published: (2025)
by: Saji, Alan, et al.
Published: (2025)
Detecting Check-Worthy Claims in Political Debates, Speeches, and Interviews Using Audio Data
by: Ivanov, Petar, et al.
Published: (2023)
by: Ivanov, Petar, et al.
Published: (2023)
ObfusQAte: A Proposed Framework to Evaluate LLM Robustness on Obfuscated Factual Question Answering
by: Ghosh, Shubhra, et al.
Published: (2025)
by: Ghosh, Shubhra, et al.
Published: (2025)
Atomic Inference for NLI with Generated Facts as Atoms
by: Stacey, Joe, et al.
Published: (2023)
by: Stacey, Joe, et al.
Published: (2023)
How Do Large Language Models Acquire Factual Knowledge During Pretraining?
by: Chang, Hoyeon, et al.
Published: (2024)
by: Chang, Hoyeon, et al.
Published: (2024)
Cultural Benchmarking of LLMs in Standard and Dialectal Arabic Dialogues
by: Kautsar, Muhammad Dehan Al, et al.
Published: (2026)
by: Kautsar, Muhammad Dehan Al, et al.
Published: (2026)
FIRE: Fact-checking with Iterative Retrieval and Verification
by: Xie, Zhuohan, et al.
Published: (2024)
by: Xie, Zhuohan, et al.
Published: (2024)
ALAS: A Stateful Multi-LLM Agent Framework for Disruption-Aware Planning
by: Chang, Edward Y., et al.
Published: (2025)
by: Chang, Edward Y., et al.
Published: (2025)
Understanding LLM Evaluator Behavior: A Structured Multi-Evaluator Framework for Merchant Risk Assessment
by: Wang, Liang, et al.
Published: (2026)
by: Wang, Liang, et al.
Published: (2026)
Evaluating Relational Reasoning in LLMs with REL
by: Fesser, Lukas, et al.
Published: (2026)
by: Fesser, Lukas, et al.
Published: (2026)
Evaluating Input Feature Explanations through a Unified Diagnostic Evaluation Framework
by: Sun, Jingyi, et al.
Published: (2024)
by: Sun, Jingyi, et al.
Published: (2024)
NeuralNexus at BEA 2025 Shared Task: Retrieval-Augmented Prompting for Mistake Identification in AI Tutors
by: Naeem, Numaan, et al.
Published: (2025)
by: Naeem, Numaan, et al.
Published: (2025)
Whose Facts Win? LLM Source Preferences under Knowledge Conflicts
by: Schuster, Jakob, et al.
Published: (2026)
by: Schuster, Jakob, et al.
Published: (2026)
ToolGen: Unified Tool Retrieval and Calling via Generation
by: Wang, Renxi, et al.
Published: (2024)
by: Wang, Renxi, et al.
Published: (2024)
An Investigation of Neuron Activation as a Unified Lens to Explain Chain-of-Thought Eliciting Arithmetic Reasoning of LLMs
by: Rai, Daking, et al.
Published: (2024)
by: Rai, Daking, et al.
Published: (2024)
Pareto-Optimized Open-Source LLMs for Healthcare via Context Retrieval
by: Bayarri-Planas, Jordi, et al.
Published: (2024)
by: Bayarri-Planas, Jordi, et al.
Published: (2024)
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
by: Ashuach, Tomer, et al.
Published: (2025)
by: Ashuach, Tomer, et al.
Published: (2025)
LLMs Are Not Scorers: Rethinking MT Evaluation with Generation-Based Methods
by: Cui, Hyang
Published: (2025)
by: Cui, Hyang
Published: (2025)
Toward Architecture-Aware Evaluation Metrics for LLM Agents
by: Souza, Débora, et al.
Published: (2026)
by: Souza, Débora, et al.
Published: (2026)
Cetvel: A Unified Benchmark for Evaluating Language Understanding, Generation and Cultural Capacity of LLMs for Turkish
by: Er, Yakup Abrek, et al.
Published: (2025)
by: Er, Yakup Abrek, et al.
Published: (2025)
From Facts to Folklore: Evaluating Large Language Models on Bengali Cultural Knowledge
by: Chowdhury, Nafis, et al.
Published: (2025)
by: Chowdhury, Nafis, et al.
Published: (2025)
ManagerBench: Evaluating the Safety-Pragmatism Trade-off in Autonomous LLMs
by: Simhi, Adi, et al.
Published: (2025)
by: Simhi, Adi, et al.
Published: (2025)
Pipeline and Dataset Generation for Automated Fact-checking in Almost Any Language
by: Drchal, Jan, et al.
Published: (2023)
by: Drchal, Jan, et al.
Published: (2023)
ReFactX: Scalable Reasoning with Reliable Facts via Constrained Generation
by: Pozzi, Riccardo, et al.
Published: (2025)
by: Pozzi, Riccardo, et al.
Published: (2025)
BOUQuET: dataset, Benchmark and Open initiative for Universal Quality Evaluation in Translation
by: The Omnilingual MT Team, et al.
Published: (2025)
by: The Omnilingual MT Team, et al.
Published: (2025)
Evaluating the Clinical Safety of LLMs in Response to High-Risk Mental Health Disclosures
by: Shah, Siddharth, et al.
Published: (2025)
by: Shah, Siddharth, et al.
Published: (2025)
Overview of the Sensemaking Task at the ELOQUENT 2025 Lab: LLMs as Teachers, Students and Evaluators
by: Šindelář, Pavel, et al.
Published: (2025)
by: Šindelář, Pavel, et al.
Published: (2025)
MemeLens: Multilingual Multitask VLMs for Memes
by: Shahroor, Ali Ezzat, et al.
Published: (2026)
by: Shahroor, Ali Ezzat, et al.
Published: (2026)
Y-NQ: English-Yorùbá Evaluation dataset for Open-Book Reading Comprehension and Text Generation
by: Costa-jussà, Marta R., et al.
Published: (2024)
by: Costa-jussà, Marta R., et al.
Published: (2024)
How LLMs Are Persuaded: A Few Attention Heads, Rerouted
by: Sun, Xiangkun, et al.
Published: (2026)
by: Sun, Xiangkun, et al.
Published: (2026)
Similar Items
-
OpenFactCheck: Building, Benchmarking Customized Fact-Checking Systems and Evaluating the Factuality of Claims and LLMs
by: Wang, Yuxia, et al.
Published: (2024) -
UrduFactCheck: An Agentic Fact-Checking Framework for Urdu with Evidence Boosting and Benchmarking
by: Ahmad, Sarfraz, et al.
Published: (2025) -
The CLEF-2025 CheckThat! Lab: Subjectivity, Fact-Checking, Claim Normalization, and Retrieval
by: Alam, Firoj, et al.
Published: (2025) -
Retrieve-Refine-Calibrate: A Framework for Complex Claim Fact-Checking
by: Sun, Mingwei, et al.
Published: (2026) -
Integrating Causal Reasoning into Automated Fact-Checking
by: Rebboud, Youssra, et al.
Published: (2025)