Evaluating open-source Large Language Models for automated fact-checking
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Fontana, Nicolo', Corso, Francesco, Zuccolotto, Enrico, Pierri, Francesco |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Do Androids Dream of Unseen Puppeteers? Probing for a Conspiracy Mindset in Large Language Models
par: Corso, Francesco, et autres
Publié: (2025)
par: Corso, Francesco, et autres
Publié: (2025)
Evaluating AI capabilities in detecting conspiracy theories on YouTube
par: La Rocca, Leonardo, et autres
Publié: (2025)
par: La Rocca, Leonardo, et autres
Publié: (2025)
Nicer Than Humans: How do Large Language Models Behave in the Prisoner's Dilemma?
par: Fontana, Nicoló, et autres
Publié: (2024)
par: Fontana, Nicoló, et autres
Publié: (2024)
Analyzing the Safety of Japanese Large Language Models in Stereotype-Triggering Prompts
par: Nakanishi, Akito, et autres
Publié: (2025)
par: Nakanishi, Akito, et autres
Publié: (2025)
Towards an Automated Framework to Audit Youth Safety on TikTok
par: Xue, Linda, et autres
Publié: (2025)
par: Xue, Linda, et autres
Publié: (2025)
A Longitudinal Study of Italian and French Reddit Conversations Around the Russian Invasion of Ukraine
par: Corso, Francesco, et autres
Publié: (2024)
par: Corso, Francesco, et autres
Publié: (2024)
Evaluating Reliability Asymmetries in Chinese Factual Search and AI Answers
par: Liu, Geng, et autres
Publié: (2025)
par: Liu, Geng, et autres
Publié: (2025)
Among Us: Language of Conspiracy Theorists on Mainstream Reddit
par: Corso, Francesco, et autres
Publié: (2025)
par: Corso, Francesco, et autres
Publié: (2025)
Effects of Algorithmic Visibility on Conspiracy Communities: Reddit after Epstein's 'Suicide'
par: Attanasio, Asja, et autres
Publié: (2025)
par: Attanasio, Asja, et autres
Publié: (2025)
Conspiracy theories and where to find them on TikTok
par: Corso, Francesco, et autres
Publié: (2024)
par: Corso, Francesco, et autres
Publié: (2024)
What we can learn from TikTok through its Research API
par: Corso, Francesco, et autres
Publié: (2024)
par: Corso, Francesco, et autres
Publié: (2024)
The Perils & Promises of Fact-checking with Large Language Models
par: Quelle, Dorian, et autres
Publié: (2023)
par: Quelle, Dorian, et autres
Publié: (2023)
Truth Sleuth and Trend Bender: AI Agents to fact-check YouTube videos and influence opinions
par: Logé, Cécile, et autres
Publié: (2025)
par: Logé, Cécile, et autres
Publié: (2025)
Lost in translation: using global fact-checks to measure multilingual misinformation prevalence, spread, and evolution
par: Quelle, Dorian, et autres
Publié: (2023)
par: Quelle, Dorian, et autres
Publié: (2023)
From Speech to Subtitles: Evaluating ASR Models in Subtitling Italian Television Programs
par: Lucca, Alessandro, et autres
Publié: (2025)
par: Lucca, Alessandro, et autres
Publié: (2025)
Overreliance on AI in Information-seeking from Video Content
par: Møller, Anders Giovanni, et autres
Publié: (2026)
par: Møller, Anders Giovanni, et autres
Publié: (2026)
Evaluating Chinese Large Language Models: The Influence of Persona Assignment on Stereotypes and Safeguards
par: Liu, Geng, et autres
Publié: (2025)
par: Liu, Geng, et autres
Publié: (2025)
Evaluating Proactive Risk Awareness of Large Language Models
par: Luo, Xuan, et autres
Publié: (2026)
par: Luo, Xuan, et autres
Publié: (2026)
Extrinsic Evaluation of Cultural Competence in Large Language Models
par: Bhatt, Shaily, et autres
Publié: (2024)
par: Bhatt, Shaily, et autres
Publié: (2024)
From Content to Audience: A Multimodal Annotation Framework for Broadcast Television Analytics
par: Cupini, Paolo, et autres
Publié: (2026)
par: Cupini, Paolo, et autres
Publié: (2026)
Towards Grammatical Tagging for the Legal Language of Cybersecurity
par: Castiglione, Gianpietro, et autres
Publié: (2023)
par: Castiglione, Gianpietro, et autres
Publié: (2023)
Anticipating Innovation Using Large Language Models
par: Fenoaltea, Enrico Maria, et autres
Publié: (2026)
par: Fenoaltea, Enrico Maria, et autres
Publié: (2026)
Speciesism in AI: Evaluating Discrimination Against Animals in Large Language Models
par: Jotautaitė, Monika, et autres
Publié: (2025)
par: Jotautaitė, Monika, et autres
Publié: (2025)
AIC CTU system at AVeriTeC: Re-framing automated fact-checking as a simple RAG task
par: Ullrich, Herbert, et autres
Publié: (2024)
par: Ullrich, Herbert, et autres
Publié: (2024)
Leveraging Large Language Models for Preliminary Security Risk Analysis: A Mission-Critical Case Study
par: Esposito, Matteo, et autres
Publié: (2024)
par: Esposito, Matteo, et autres
Publié: (2024)
Cancer-Myth: Evaluating Large Language Models on Patient Questions with False Presuppositions
par: Zhu, Wang Bill, et autres
Publié: (2025)
par: Zhu, Wang Bill, et autres
Publié: (2025)
Not All Jokes Land: Evaluating Large Language Models Understanding of Workplace Humor
par: Shafiei, Mohammadamin, et autres
Publié: (2025)
par: Shafiei, Mohammadamin, et autres
Publié: (2025)
Leveraging Large Language Models for Actionable Course Evaluation Student Feedback to Lecturers
par: Zhang, Mike, et autres
Publié: (2024)
par: Zhang, Mike, et autres
Publié: (2024)
Optimizing Large Language Models for ESG Activity Detection in Financial Texts
par: Birti, Mattia, et autres
Publié: (2025)
par: Birti, Mattia, et autres
Publié: (2025)
The doctor will polygraph you now: ethical concerns with AI for fact-checking patients
par: Anibal, James, et autres
Publié: (2024)
par: Anibal, James, et autres
Publié: (2024)
Evaluating Large Language Models for Detecting Antisemitism
par: Patel, Jay, et autres
Publié: (2025)
par: Patel, Jay, et autres
Publié: (2025)
Evaluating Large Language Models on Spatial Tasks: A Multi-Task Benchmarking Study
par: Xu, Liuchang, et autres
Publié: (2024)
par: Xu, Liuchang, et autres
Publié: (2024)
Evaluating Psychological Safety of Large Language Models
par: Li, Xingxuan, et autres
Publié: (2022)
par: Li, Xingxuan, et autres
Publié: (2022)
Evaluating Large Language Models in Theory of Mind Tasks
par: Kosinski, Michal
Publié: (2023)
par: Kosinski, Michal
Publié: (2023)
QuanTemp: A real-world open-domain benchmark for fact-checking numerical claims
par: V, Venktesh, et autres
Publié: (2024)
par: V, Venktesh, et autres
Publié: (2024)
Berta: an open-source, modular tool for AI-enabled clinical documentation
par: Vaid, Samridhi, et autres
Publié: (2026)
par: Vaid, Samridhi, et autres
Publié: (2026)
Motivation in Large Language Models
par: Nahum, Omer, et autres
Publié: (2026)
par: Nahum, Omer, et autres
Publié: (2026)
Comparing diversity, negativity, and stereotypes in Chinese-language AI technologies: an investigation of Baidu, Ernie and Qwen
par: Liu, Geng, et autres
Publié: (2024)
par: Liu, Geng, et autres
Publié: (2024)
Human vs. Machine: Behavioral Differences Between Expert Humans and Language Models in Wargame Simulations
par: Lamparth, Max, et autres
Publié: (2024)
par: Lamparth, Max, et autres
Publié: (2024)
From Trust to Truth: Actionable policies for the use of AI in fact-checking in Germany and Ukraine
par: Solopova, Veronika
Publié: (2025)
par: Solopova, Veronika
Publié: (2025)
Documents similaires
-
Do Androids Dream of Unseen Puppeteers? Probing for a Conspiracy Mindset in Large Language Models
par: Corso, Francesco, et autres
Publié: (2025) -
Evaluating AI capabilities in detecting conspiracy theories on YouTube
par: La Rocca, Leonardo, et autres
Publié: (2025) -
Nicer Than Humans: How do Large Language Models Behave in the Prisoner's Dilemma?
par: Fontana, Nicoló, et autres
Publié: (2024) -
Analyzing the Safety of Japanese Large Language Models in Stereotype-Triggering Prompts
par: Nakanishi, Akito, et autres
Publié: (2025) -
Towards an Automated Framework to Audit Youth Safety on TikTok
par: Xue, Linda, et autres
Publié: (2025)