Loki: An Open-Source Tool for Fact Verification
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Haonan, Han, Xudong, Wang, Hao, Wang, Yuxia, Wang, Minghan, Xing, Rui, Geng, Yilin, Zhai, Zenan, Nakov, Preslav, Baldwin, Timothy |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
OpenFactCheck: Building, Benchmarking Customized Fact-Checking Systems and Evaluating the Factuality of Claims and LLMs
di: Wang, Yuxia, et al.
Pubblicazione: (2024)
di: Wang, Yuxia, et al.
Pubblicazione: (2024)
A Chinese Dataset for Evaluating the Safeguards in Large Language Models
di: Wang, Yuxia, et al.
Pubblicazione: (2024)
di: Wang, Yuxia, et al.
Pubblicazione: (2024)
Rethinking STS and NLI in Large Language Models
di: Wang, Yuxia, et al.
Pubblicazione: (2023)
di: Wang, Yuxia, et al.
Pubblicazione: (2023)
OpenFactCheck: A Unified Framework for Factuality Evaluation of LLMs
di: Iqbal, Hasan, et al.
Pubblicazione: (2024)
di: Iqbal, Hasan, et al.
Pubblicazione: (2024)
COMMUNITYNOTES: A Dataset for Exploring the Helpfulness of Fact-Checking Explanations
di: Xing, Rui, et al.
Pubblicazione: (2025)
di: Xing, Rui, et al.
Pubblicazione: (2025)
FIRE: Fact-checking with Iterative Retrieval and Verification
di: Xie, Zhuohan, et al.
Pubblicazione: (2024)
di: Xie, Zhuohan, et al.
Pubblicazione: (2024)
Can Machines Resonate with Humans? Evaluating the Emotional and Empathic Comprehension of LMs
di: Manzoor, Muhammad Arslan, et al.
Pubblicazione: (2024)
di: Manzoor, Muhammad Arslan, et al.
Pubblicazione: (2024)
Arabic Dataset for LLM Safeguard Evaluation
di: Ashraf, Yasser, et al.
Pubblicazione: (2024)
di: Ashraf, Yasser, et al.
Pubblicazione: (2024)
Against The Achilles' Heel: A Survey on Red Teaming for Generative Models
di: Lin, Lizhi, et al.
Pubblicazione: (2024)
di: Lin, Lizhi, et al.
Pubblicazione: (2024)
RuozhiBench: Evaluating LLMs with Logical Fallacies and Misleading Premises
di: Zhai, Zenan, et al.
Pubblicazione: (2025)
di: Zhai, Zenan, et al.
Pubblicazione: (2025)
UnsafeChain: Enhancing Reasoning Model Safety via Hard Cases
di: Tomar, Raj Vardhan, et al.
Pubblicazione: (2025)
di: Tomar, Raj Vardhan, et al.
Pubblicazione: (2025)
How Does Prefix Matter in Reasoning Model Tuning?
di: Tomar, Raj Vardhan, et al.
Pubblicazione: (2026)
di: Tomar, Raj Vardhan, et al.
Pubblicazione: (2026)
ToolGen: Unified Tool Retrieval and Calling via Generation
di: Wang, Renxi, et al.
Pubblicazione: (2024)
di: Wang, Renxi, et al.
Pubblicazione: (2024)
MuDRiC: Multi-Dialect Reasoning for Arabic Commonsense Validation
di: Elozeiri, Kareem, et al.
Pubblicazione: (2025)
di: Elozeiri, Kareem, et al.
Pubblicazione: (2025)
Multimodal Large Language Models to Support Real-World Fact-Checking
di: Geng, Jiahui, et al.
Pubblicazione: (2024)
di: Geng, Jiahui, et al.
Pubblicazione: (2024)
TART: An Open-Source Tool-Augmented Framework for Explainable Table-based Reasoning
di: Lu, Xinyuan, et al.
Pubblicazione: (2024)
di: Lu, Xinyuan, et al.
Pubblicazione: (2024)
Demystifying Instruction Mixing for Fine-tuning Large Language Models
di: Wang, Renxi, et al.
Pubblicazione: (2023)
di: Wang, Renxi, et al.
Pubblicazione: (2023)
An Analytical Emotion Framework of Rumour Threads on Social Media
di: Xing, Rui, et al.
Pubblicazione: (2025)
di: Xing, Rui, et al.
Pubblicazione: (2025)
Evaluating Evidence Attribution in Generated Fact Checking Explanations
di: Xing, Rui, et al.
Pubblicazione: (2024)
di: Xing, Rui, et al.
Pubblicazione: (2024)
Factuality of Large Language Models: A Survey
di: Wang, Yuxia, et al.
Pubblicazione: (2024)
di: Wang, Yuxia, et al.
Pubblicazione: (2024)
HALF: Harm-Aware LLM Fairness Evaluation Aligned with Deployment
di: Mekky, Ali, et al.
Pubblicazione: (2025)
di: Mekky, Ali, et al.
Pubblicazione: (2025)
A Survey of Confidence Estimation and Calibration in Large Language Models
di: Geng, Jiahui, et al.
Pubblicazione: (2023)
di: Geng, Jiahui, et al.
Pubblicazione: (2023)
Co-FactChecker: A Framework for Human-AI Collaborative Claim Verification Using Large Reasoning Models
di: Sahnan, Dhruv, et al.
Pubblicazione: (2026)
di: Sahnan, Dhruv, et al.
Pubblicazione: (2026)
From Chaos to Clarity: Claim Normalization to Empower Fact-Checking
di: Sundriyal, Megha, et al.
Pubblicazione: (2023)
di: Sundriyal, Megha, et al.
Pubblicazione: (2023)
UrduFactCheck: An Agentic Fact-Checking Framework for Urdu with Evidence Boosting and Benchmarking
di: Ahmad, Sarfraz, et al.
Pubblicazione: (2025)
di: Ahmad, Sarfraz, et al.
Pubblicazione: (2025)
Learning From Failure: Integrating Negative Examples when Fine-tuning Large Language Models as Agents
di: Wang, Renxi, et al.
Pubblicazione: (2024)
di: Wang, Renxi, et al.
Pubblicazione: (2024)
Control Illusion: The Failure of Instruction Hierarchies in Large Language Models
di: Geng, Yilin, et al.
Pubblicazione: (2025)
di: Geng, Yilin, et al.
Pubblicazione: (2025)
Fact-Checking the Output of Large Language Models via Token-Level Uncertainty Quantification
di: Fadeeva, Ekaterina, et al.
Pubblicazione: (2024)
di: Fadeeva, Ekaterina, et al.
Pubblicazione: (2024)
Faithfulness-Aware Uncertainty Quantification for Fact-Checking the Output of Retrieval Augmented Generation
di: Fadeeva, Ekaterina, et al.
Pubblicazione: (2025)
di: Fadeeva, Ekaterina, et al.
Pubblicazione: (2025)
Qorgau: Evaluating LLM Safety in Kazakh-Russian Bilingual Contexts
di: Goloburda, Maiya, et al.
Pubblicazione: (2025)
di: Goloburda, Maiya, et al.
Pubblicazione: (2025)
Cross-Cultural Transfer of Commonsense Reasoning in LLMs: Evidence from the Arab World
di: Almheiri, Saeed, et al.
Pubblicazione: (2025)
di: Almheiri, Saeed, et al.
Pubblicazione: (2025)
Generating Zero-shot Abstractive Explanations for Rumour Verification
di: Bilal, Iman Munire, et al.
Pubblicazione: (2024)
di: Bilal, Iman Munire, et al.
Pubblicazione: (2024)
Factcheck-Bench: Fine-Grained Evaluation Benchmark for Automatic Fact-checkers
di: Wang, Yuxia, et al.
Pubblicazione: (2023)
di: Wang, Yuxia, et al.
Pubblicazione: (2023)
From Leaky Thoughts to Private Reasoning: Controlling What LRMs Say to Themselves
di: Puerto, Haritz, et al.
Pubblicazione: (2026)
di: Puerto, Haritz, et al.
Pubblicazione: (2026)
Instruction Tuning on Public Government and Cultural Data for Low-Resource Language: a Case Study in Kazakh
di: Laiyk, Nurkhan, et al.
Pubblicazione: (2025)
di: Laiyk, Nurkhan, et al.
Pubblicazione: (2025)
Atomic Reasoning for Scientific Table Claim Verification
di: Zhang, Yuji, et al.
Pubblicazione: (2025)
di: Zhang, Yuji, et al.
Pubblicazione: (2025)
Profiling News Media for Factuality and Bias Using LLMs and the Fact-Checking Methodology of Human Experts
di: Mujahid, Zain Muhammad, et al.
Pubblicazione: (2025)
di: Mujahid, Zain Muhammad, et al.
Pubblicazione: (2025)
Unconditional Truthfulness: Learning Unconditional Uncertainty of Large Language Models
di: Vazhentsev, Artem, et al.
Pubblicazione: (2024)
di: Vazhentsev, Artem, et al.
Pubblicazione: (2024)
A Comprehensive Survey of Machine Unlearning Techniques for Large Language Models
di: Geng, Jiahui, et al.
Pubblicazione: (2025)
di: Geng, Jiahui, et al.
Pubblicazione: (2025)
SCALAR: Scientific Citation-based Live Assessment of Long-context Academic Reasoning
di: Wang, Renxi, et al.
Pubblicazione: (2025)
di: Wang, Renxi, et al.
Pubblicazione: (2025)
Documenti analoghi
-
OpenFactCheck: Building, Benchmarking Customized Fact-Checking Systems and Evaluating the Factuality of Claims and LLMs
di: Wang, Yuxia, et al.
Pubblicazione: (2024) -
A Chinese Dataset for Evaluating the Safeguards in Large Language Models
di: Wang, Yuxia, et al.
Pubblicazione: (2024) -
Rethinking STS and NLI in Large Language Models
di: Wang, Yuxia, et al.
Pubblicazione: (2023) -
OpenFactCheck: A Unified Framework for Factuality Evaluation of LLMs
di: Iqbal, Hasan, et al.
Pubblicazione: (2024) -
COMMUNITYNOTES: A Dataset for Exploring the Helpfulness of Fact-Checking Explanations
di: Xing, Rui, et al.
Pubblicazione: (2025)