Factcheck-Bench: Fine-Grained Evaluation Benchmark for Automatic Fact-checkers
Fuente:
arXiv
Guardado en:
| Autores principales: | Wang, Yuxia, Reddy, Revanth Gangi, Mujahid, Zain Muhammad, Arora, Arnav, Rubashevskii, Aleksandr, Geng, Jiahui, Afzal, Osama Mohammed, Pan, Liangming, Borenstein, Nadav, Pillai, Aditya, Augenstein, Isabelle, Gurevych, Iryna, Nakov, Preslav |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Investigating Human Values in Online Communities
por: Borenstein, Nadav, et al.
Publicado: (2024)
por: Borenstein, Nadav, et al.
Publicado: (2024)
Multimodal Large Language Models to Support Real-World Fact-Checking
por: Geng, Jiahui, et al.
Publicado: (2024)
por: Geng, Jiahui, et al.
Publicado: (2024)
Beyond "Not Novel Enough": Enriching Scholarly Critique with LLM-Assisted Feedback
por: Afzal, Osama Mohammed, et al.
Publicado: (2025)
por: Afzal, Osama Mohammed, et al.
Publicado: (2025)
Can Community Notes Replace Professional Fact-Checkers?
por: Borenstein, Nadav, et al.
Publicado: (2025)
por: Borenstein, Nadav, et al.
Publicado: (2025)
OpenFactCheck: A Unified Framework for Factuality Evaluation of LLMs
por: Iqbal, Hasan, et al.
Publicado: (2024)
por: Iqbal, Hasan, et al.
Publicado: (2024)
FIRE: Fact-checking with Iterative Retrieval and Verification
por: Xie, Zhuohan, et al.
Publicado: (2024)
por: Xie, Zhuohan, et al.
Publicado: (2024)
Revealing Fine-Grained Values and Opinions in Large Language Models
por: Wright, Dustin, et al.
Publicado: (2024)
por: Wright, Dustin, et al.
Publicado: (2024)
A Survey of Confidence Estimation and Calibration in Large Language Models
por: Geng, Jiahui, et al.
Publicado: (2023)
por: Geng, Jiahui, et al.
Publicado: (2023)
Adaptive Conformal Prediction for Improving Factuality of Generations by Large Language Models
por: Rubashevskii, Aleksandr, et al.
Publicado: (2026)
por: Rubashevskii, Aleksandr, et al.
Publicado: (2026)
Profiling News Media for Factuality and Bias Using LLMs and the Fact-Checking Methodology of Human Experts
por: Mujahid, Zain Muhammad, et al.
Publicado: (2025)
por: Mujahid, Zain Muhammad, et al.
Publicado: (2025)
BiasGym: A Simple and Generalizable Framework for Analyzing and Removing Biases through Elicitation
por: Islam, Sekh Mainul, et al.
Publicado: (2025)
por: Islam, Sekh Mainul, et al.
Publicado: (2025)
Con Instruction: Universal Jailbreaking of Multimodal Large Language Models via Non-Textual Modalities
por: Geng, Jiahui, et al.
Publicado: (2025)
por: Geng, Jiahui, et al.
Publicado: (2025)
Co-FactChecker: A Framework for Human-AI Collaborative Claim Verification Using Large Reasoning Models
por: Sahnan, Dhruv, et al.
Publicado: (2026)
por: Sahnan, Dhruv, et al.
Publicado: (2026)
Stress Testing Factual Consistency Metrics for Long-Document Summarization
por: Mujahid, Zain Muhammad, et al.
Publicado: (2025)
por: Mujahid, Zain Muhammad, et al.
Publicado: (2025)
AICD Bench: A Challenging Benchmark for AI-Generated Code Detection
por: Orel, Daniil, et al.
Publicado: (2026)
por: Orel, Daniil, et al.
Publicado: (2026)
OpenFactCheck: Building, Benchmarking Customized Fact-Checking Systems and Evaluating the Factuality of Claims and LLMs
por: Wang, Yuxia, et al.
Publicado: (2024)
por: Wang, Yuxia, et al.
Publicado: (2024)
ConspirED: A Dataset for Cognitive Traits of Conspiracy Theories and Large Language Model Safety
por: Bates, Luke, et al.
Publicado: (2025)
por: Bates, Luke, et al.
Publicado: (2025)
Missci: Reconstructing Fallacies in Misrepresented Science
por: Glockner, Max, et al.
Publicado: (2024)
por: Glockner, Max, et al.
Publicado: (2024)
$\texttt{Droid}$: A Resource Suite for AI-Generated Code Detection
por: Orel, Daniil, et al.
Publicado: (2025)
por: Orel, Daniil, et al.
Publicado: (2025)
Grounding Fallacies Misrepresenting Scientific Publications in Evidence
por: Glockner, Max, et al.
Publicado: (2024)
por: Glockner, Max, et al.
Publicado: (2024)
Can Transformers Learn $n$-gram Language Models?
por: Svete, Anej, et al.
Publicado: (2024)
por: Svete, Anej, et al.
Publicado: (2024)
A Template Is All You Meme
por: Bates, Luke, et al.
Publicado: (2023)
por: Bates, Luke, et al.
Publicado: (2023)
Probing Pre-Trained Language Models for Cross-Cultural Differences in Values
por: Arora, Arnav, et al.
Publicado: (2022)
por: Arora, Arnav, et al.
Publicado: (2022)
Why Should This Article Be Deleted? Transparent Stance Detection in Multilingual Wikipedia Editor Discussions
por: Kaffee, Lucie-Aimée, et al.
Publicado: (2023)
por: Kaffee, Lucie-Aimée, et al.
Publicado: (2023)
Revisiting Noise in Natural Language Processing for Computational Social Science
por: Borenstein, Nadav
Publicado: (2025)
por: Borenstein, Nadav
Publicado: (2025)
M4FC: a Multimodal, Multilingual, Multicultural, Multitask Real-World Fact-Checking Dataset
por: Geng, Jiahui, et al.
Publicado: (2025)
por: Geng, Jiahui, et al.
Publicado: (2025)
Faithfulness-Aware Uncertainty Quantification for Fact-Checking the Output of Retrieval Augmented Generation
por: Fadeeva, Ekaterina, et al.
Publicado: (2025)
por: Fadeeva, Ekaterina, et al.
Publicado: (2025)
Uncertainty Quantification for LLMs through Minimum Bayes Risk: Bridging Confidence and Consistency
por: Vashurin, Roman, et al.
Publicado: (2025)
por: Vashurin, Roman, et al.
Publicado: (2025)
Community Moderation and the New Epistemology of Fact Checking on Social Media
por: Augenstein, Isabelle, et al.
Publicado: (2025)
por: Augenstein, Isabelle, et al.
Publicado: (2025)
Rethinking STS and NLI in Large Language Models
por: Wang, Yuxia, et al.
Publicado: (2023)
por: Wang, Yuxia, et al.
Publicado: (2023)
Can LLMs Automate Fact-Checking Article Writing?
por: Sahnan, Dhruv, et al.
Publicado: (2025)
por: Sahnan, Dhruv, et al.
Publicado: (2025)
Unstructured Evidence Attribution for Long Context Query Focused Summarization
por: Wright, Dustin, et al.
Publicado: (2025)
por: Wright, Dustin, et al.
Publicado: (2025)
Presumed Cultural Identity: How Names Shape LLM Responses
por: Pawar, Siddhesh, et al.
Publicado: (2025)
por: Pawar, Siddhesh, et al.
Publicado: (2025)
Fact-Checking the Output of Large Language Models via Token-Level Uncertainty Quantification
por: Fadeeva, Ekaterina, et al.
Publicado: (2024)
por: Fadeeva, Ekaterina, et al.
Publicado: (2024)
From Chaos to Clarity: Claim Normalization to Empower Fact-Checking
por: Sundriyal, Megha, et al.
Publicado: (2023)
por: Sundriyal, Megha, et al.
Publicado: (2023)
UnsafeChain: Enhancing Reasoning Model Safety via Hard Cases
por: Tomar, Raj Vardhan, et al.
Publicado: (2025)
por: Tomar, Raj Vardhan, et al.
Publicado: (2025)
How Does Prefix Matter in Reasoning Model Tuning?
por: Tomar, Raj Vardhan, et al.
Publicado: (2026)
por: Tomar, Raj Vardhan, et al.
Publicado: (2026)
Don't Throw Away Your Beams: Improving Consistency-based Uncertainties in LLMs via Beam Search
por: Fadeeva, Ekaterina, et al.
Publicado: (2025)
por: Fadeeva, Ekaterina, et al.
Publicado: (2025)
Multi-Modal Framing Analysis of News
por: Arora, Arnav, et al.
Publicado: (2025)
por: Arora, Arnav, et al.
Publicado: (2025)
Show Me the Work: Fact-Checkers' Requirements for Explainable Automated Fact-Checking
por: Warren, Greta, et al.
Publicado: (2025)
por: Warren, Greta, et al.
Publicado: (2025)
Ejemplares similares
-
Investigating Human Values in Online Communities
por: Borenstein, Nadav, et al.
Publicado: (2024) -
Multimodal Large Language Models to Support Real-World Fact-Checking
por: Geng, Jiahui, et al.
Publicado: (2024) -
Beyond "Not Novel Enough": Enriching Scholarly Critique with LLM-Assisted Feedback
por: Afzal, Osama Mohammed, et al.
Publicado: (2025) -
Can Community Notes Replace Professional Fact-Checkers?
por: Borenstein, Nadav, et al.
Publicado: (2025) -
OpenFactCheck: A Unified Framework for Factuality Evaluation of LLMs
por: Iqbal, Hasan, et al.
Publicado: (2024)