MAFALDA: A Benchmark and Comprehensive Study of Fallacy Detection and Classification
Fuente:
arXiv
Saved in:
| Main Authors: | Helwe, Chadi, Calamai, Tom, Paris, Pierre-Henri, Clavel, Chloé, Suchanek, Fabian |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SmolKalam: Ensemble Quality-Filtered Translation at Scale for High Quality Arabic Post-Training Data
by: Alrashed, Sultan, et al.
Published: (2025)
by: Alrashed, Sultan, et al.
Published: (2025)
Graphically Speaking: Unmasking Abuse in Social Media with Conversation Insights
by: Nouri, Célia, et al.
Published: (2025)
by: Nouri, Célia, et al.
Published: (2025)
MisSynth: Improving MISSCI Logical Fallacies Classification with Synthetic Data
by: Poliakov, Mykhailo, et al.
Published: (2025)
by: Poliakov, Mykhailo, et al.
Published: (2025)
A Logical Fallacy-Informed Framework for Argument Generation
by: Mouchel, Luca, et al.
Published: (2024)
by: Mouchel, Luca, et al.
Published: (2024)
Autoformalizing Natural Language to First-Order Logic: A Case Study in Logical Fallacy Detection
by: Lalwani, Abhinav, et al.
Published: (2024)
by: Lalwani, Abhinav, et al.
Published: (2024)
The Factuality of Large Language Models in the Legal Domain
by: Hamdani, Rajaa El, et al.
Published: (2024)
by: Hamdani, Rajaa El, et al.
Published: (2024)
Can We Count on LLMs? The Fixed-Effect Fallacy and Claims of GPT-4 Capabilities
by: Ball, Thomas, et al.
Published: (2024)
by: Ball, Thomas, et al.
Published: (2024)
MedHallu: A Comprehensive Benchmark for Detecting Medical Hallucinations in Large Language Models
by: Pandit, Shrey, et al.
Published: (2025)
by: Pandit, Shrey, et al.
Published: (2025)
Addressing the Ecological Fallacy in Larger LMs with Human Context
by: Soni, Nikita, et al.
Published: (2026)
by: Soni, Nikita, et al.
Published: (2026)
Large Language Models Are Better Logical Fallacy Reasoners with Counterargument, Explanation, and Goal-Aware Prompt Formulation
by: Jeong, Jiwon, et al.
Published: (2025)
by: Jeong, Jiwon, et al.
Published: (2025)
Hierarchical Multi-Label Classification of Online Vaccine Concerns
by: Zhu, Chloe Qinyu, et al.
Published: (2024)
by: Zhu, Chloe Qinyu, et al.
Published: (2024)
Benchmarking ChatGPT on Algorithmic Reasoning
by: McLeish, Sean, et al.
Published: (2024)
by: McLeish, Sean, et al.
Published: (2024)
Fine-Grained Benchmark Generation for Comprehensive Evaluation of Foundation Models
by: Islam, Mohammed Saidul, et al.
Published: (2026)
by: Islam, Mohammed Saidul, et al.
Published: (2026)
Detecting Greenwashing: A Natural Language Processing Literature Survey
by: Calamai, Tom, et al.
Published: (2025)
by: Calamai, Tom, et al.
Published: (2025)
Do Language Models Enjoy Their Own Stories? Prompting Large Language Models for Automatic Story Evaluation
by: Chhun, Cyril, et al.
Published: (2024)
by: Chhun, Cyril, et al.
Published: (2024)
CALM : A Multi-task Benchmark for Comprehensive Assessment of Language Model Bias
by: Gupta, Vipul, et al.
Published: (2023)
by: Gupta, Vipul, et al.
Published: (2023)
FusionBench: A Unified Library and Comprehensive Benchmark for Deep Model Fusion
by: Tang, Anke, et al.
Published: (2024)
by: Tang, Anke, et al.
Published: (2024)
Classification of User Reports for Detection of Faulty Computer Components using NLP Models: A Case Study
by: Silva, Maria de Lourdes M., et al.
Published: (2025)
by: Silva, Maria de Lourdes M., et al.
Published: (2025)
DeceptionBench: A Comprehensive Benchmark for AI Deception Behaviors in Real-world Scenarios
by: Huang, Yao, et al.
Published: (2025)
by: Huang, Yao, et al.
Published: (2025)
CTBench: A Comprehensive Benchmark for Evaluating Language Model Capabilities in Clinical Trial Design
by: Neehal, Nafis, et al.
Published: (2024)
by: Neehal, Nafis, et al.
Published: (2024)
LLMs Are Not Intelligent Thinkers: Introducing Mathematical Topic Tree Benchmark for Comprehensive Evaluation of LLMs
by: Davoodi, Arash Gholami, et al.
Published: (2024)
by: Davoodi, Arash Gholami, et al.
Published: (2024)
KFinEval-Pilot: A Comprehensive Benchmark Suite for Korean Financial Language Understanding
by: Hwang, Bokwang, et al.
Published: (2025)
by: Hwang, Bokwang, et al.
Published: (2025)
CS-Bench: A Comprehensive Benchmark for Large Language Models towards Computer Science Mastery
by: Song, Xiaoshuai, et al.
Published: (2024)
by: Song, Xiaoshuai, et al.
Published: (2024)
Log-Likelihood, Simpson's Paradox, and the Detection of Machine-Generated Text
by: Kempton, Tom, et al.
Published: (2026)
by: Kempton, Tom, et al.
Published: (2026)
A Comprehensive Study on Quantization Techniques for Large Language Models
by: Lang, Jiedong, et al.
Published: (2024)
by: Lang, Jiedong, et al.
Published: (2024)
BELLA: Black box model Explanations by Local Linear Approximations
by: Radulovic, Nedeljko, et al.
Published: (2023)
by: Radulovic, Nedeljko, et al.
Published: (2023)
Comprehensive Study on German Language Models for Clinical and Biomedical Text Understanding
by: Idrissi-Yaghir, Ahmad, et al.
Published: (2024)
by: Idrissi-Yaghir, Ahmad, et al.
Published: (2024)
DrBenchmark: A Large Language Understanding Evaluation Benchmark for French Biomedical Domain
by: Labrak, Yanis, et al.
Published: (2024)
by: Labrak, Yanis, et al.
Published: (2024)
Team QUST at SemEval-2023 Task 3: A Comprehensive Study of Monolingual and Multilingual Approaches for Detecting Online News Genre, Framing and Persuasion Techniques
by: Jiang, Ye
Published: (2023)
by: Jiang, Ye
Published: (2023)
MEDEC: A Benchmark for Medical Error Detection and Correction in Clinical Notes
by: Abacha, Asma Ben, et al.
Published: (2024)
by: Abacha, Asma Ben, et al.
Published: (2024)
Text-ADBench: Text Anomaly Detection Benchmark based on LLMs Embedding
by: Xiao, Feng, et al.
Published: (2025)
by: Xiao, Feng, et al.
Published: (2025)
AI-Generated Text Detection and Classification Based on BERT Deep Learning Algorithm
by: Wang, Hao, et al.
Published: (2024)
by: Wang, Hao, et al.
Published: (2024)
RAT-Bench: A Comprehensive Benchmark for Text Anonymization
by: Krčo, Nataša, et al.
Published: (2026)
by: Krčo, Nataša, et al.
Published: (2026)
FanChuan: A Multilingual and Graph-Structured Benchmark For Parody Detection and Analysis
by: Zheng, Yilun, et al.
Published: (2025)
by: Zheng, Yilun, et al.
Published: (2025)
Qiana: A First-Order Formalism to Quantify over Contexts and Formulas with Temporality
by: Coumes, Simon, et al.
Published: (2026)
by: Coumes, Simon, et al.
Published: (2026)
PII-Scope: A Comprehensive Study on Training Data PII Extraction Attacks in LLMs
by: Nakka, Krishna Kanth, et al.
Published: (2024)
by: Nakka, Krishna Kanth, et al.
Published: (2024)
Retrieval Augmented Generation or Long-Context LLMs? A Comprehensive Study and Hybrid Approach
by: Li, Zhuowan, et al.
Published: (2024)
by: Li, Zhuowan, et al.
Published: (2024)
Benchmarking for Domain-Specific LLMs: A Case Study on Academia and Beyond
by: Chen, Rubing, et al.
Published: (2025)
by: Chen, Rubing, et al.
Published: (2025)
Automatic Classification of User Requirements from Online Feedback -- A Replication Study
by: Bhatt, Meet, et al.
Published: (2025)
by: Bhatt, Meet, et al.
Published: (2025)
SMS Spam Detection and Classification to Combat Abuse in Telephone Networks Using Natural Language Processing
by: Oyeyemi, Dare Azeez, et al.
Published: (2024)
by: Oyeyemi, Dare Azeez, et al.
Published: (2024)
Similar Items
-
SmolKalam: Ensemble Quality-Filtered Translation at Scale for High Quality Arabic Post-Training Data
by: Alrashed, Sultan, et al.
Published: (2025) -
Graphically Speaking: Unmasking Abuse in Social Media with Conversation Insights
by: Nouri, Célia, et al.
Published: (2025) -
MisSynth: Improving MISSCI Logical Fallacies Classification with Synthetic Data
by: Poliakov, Mykhailo, et al.
Published: (2025) -
A Logical Fallacy-Informed Framework for Argument Generation
by: Mouchel, Luca, et al.
Published: (2024) -
Autoformalizing Natural Language to First-Order Logic: A Case Study in Logical Fallacy Detection
by: Lalwani, Abhinav, et al.
Published: (2024)